From Ingress to Inference Gateway: Gateway API and the Inference Extension
Traditional load balancing assumes all backends are the same, so it does not matter which one handles a request. LLM traffic changes this. For example, one vLLM replica might have a warm KV cache for
todea.hashnode.dev19 min read