The five-stage ladder is a useful way to frame this: check the quick wins (replicas, prefix-cache routing, queue behavior) before tearing into kernels or NCCL configs. Using kv-cache utilization as a fork between tuning and scaling is a detail I will remember.
The tension in stage one is worth naming directly: routing for prefix-cache locality and routing for load balance pull in opposite directions once a workload has both very short and very long sessions. If you pin a conversation to the replica that holds its KV cache prefix, a long-running heavy session keeps landing on the same pod and grows its queue exactly the way the round-robin section describes, except now it's a deliberate design choice instead of an accident. Gateway API Inference Extension and llm-d are named as the fix for the guessing load balancer, do they actually resolve that tradeoff, or do they just make the tradeoff visible and require someone to set a policy for when cache locality loses to queue depth?
Puneet Khandelwal
Getting inference latency right on K8s is half about pod autoscaling and half about stopping noisy neighbors from wrecking tail latencies. Read this if you run models in production.