The 35 GB INT4 example is a useful starting point, but the usable context length deserves its own memory budget. The compressed weights can fit while a longer conversation still runs out of memory because the attention cache and runtime buffers also occupy the same pool. That makes 'the model loads' a weaker test than 'the intended conversation runs.'
For the laptop reality check, I would record memory and response speed at several prompt lengths with the same offload settings. Leaving headroom for the operating system also matters on unified-memory machines. It would make the hardware examples easier to translate into an actual local workload.
The 35 GB INT4 example is a useful starting point, but the usable context length deserves its own memory budget. The compressed weights can fit while a longer conversation still runs out of memory because the attention cache and runtime buffers also occupy the same pool. That makes 'the model loads' a weaker test than 'the intended conversation runs.'
For the laptop reality check, I would record memory and response speed at several prompt lengths with the same offload settings. Leaving headroom for the operating system also matters on unified-memory machines. It would make the hardware examples easier to translate into an actual local workload.