The useful shift here is treating model sizing as a workload problem rather than a download-size problem. A model can technically fit once, yet become impractical as context length, KV cache, concurrency, and batching increase.
For agentic workloads, I’d add one more dimension to the checklist: memory per active request under realistic concurrency. A configuration that looks comfortable for one long-context session can behave very differently with 10–20 simultaneous sessions, because KV cache becomes a shared capacity constraint while continuous batching changes the throughput/latency trade-off.
That also makes quantization decisions less binary. The right question isn't simply “what is the smallest quantization that fits?” but “which quantization leaves enough memory headroom for the context and concurrency profile we actually need?”
The useful shift here is treating model sizing as a workload problem rather than a download-size problem. A model can technically fit once, yet become impractical as context length, KV cache, concurrency, and batching increase.
For agentic workloads, I’d add one more dimension to the checklist: memory per active request under realistic concurrency. A configuration that looks comfortable for one long-context session can behave very differently with 10–20 simultaneous sessions, because KV cache becomes a shared capacity constraint while continuous batching changes the throughput/latency trade-off.
That also makes quantization decisions less binary. The right question isn't simply “what is the smallest quantization that fits?” but “which quantization leaves enough memory headroom for the context and concurrency profile we actually need?”