Great point on memory per active request under realistic concurrency! A setup that comfortably holds 15 sessions at the start might only hold a handful an hour later. Agree on quantization too: dropping from 16-bit to 4-bit weights is about making the model fit, and at the same time about buying room for more concurrency. KV cache quantization also helps in this case.

