The tension in stage one is worth naming directly: routing for prefix-cache locality and routing for load balance pull in opposite directions once a workload has both very short and very long sessions. If you pin a conversation to the replica that holds its KV cache prefix, a long-running heavy session keeps landing on the same pod and grows its queue exactly the way the round-robin section describes, except now it's a deliberate design choice instead of an accident. Gateway API Inference Extension and llm-d are named as the fix for the guessing load balancer, do they actually resolve that tradeoff, or do they just make the tradeoff visible and require someone to set a policy for when cache locality loses to queue depth?