In modern Large Language Model (LLM) architectures, two foundational engineering bottlenecks dictate model scalability and inference latency: Representing Order: How self-attention mechanisms—which a
rahul-ai.hashnode.dev15 min readAI Engineer. Computer Vision, RAG and LLM agents.
Putting RoPE and the KV cache family in one article makes sense because they are the two things that decide what long context actually costs you. One framing worth adding for readers coming from the serving side: MQA and GQA are memory bandwidth optimisations more than compute ones. Decoding is bandwidth-bound - you re-read the whole cache every token - so shrinking the number of KV heads speeds up generation even when FLOPs barely move, which surprises people who benchmark on prefill and see nothing. The practical consequence of the cache section is that serving economics stop being about tokens per second and start being about how many contexts fit in memory at once. Long-running agent workloads keep large caches resident across many steps, so cost per resident context-hour predicts your bill better than cost per token does. Worth noting on RoPE too that extending context at inference is not free - the positions the model never trained on behave differently, which is why scaling tricks exist and why long-context benchmarks disagree with each other so much.
Ahmet Özel
@Rahul Sai Indeevar V - adding this here since the reply box on your thread would not submit for me.
If you do write that update, the experiment that makes the prefill and decode split click is measuring both phases separately at two or three batch sizes. Prefill throughput scales with batch until you saturate compute; decode throughput barely moves, because you re-read the whole KV cache either way. Seeing those two curves next to each other explains why continuous batching exists, and why a single tokens-per-second figure hides most of what is actually happening.