Running LLMs on Kubernetes in Production: KServe, vLLM, llm-d, Envoy AI Gateway, and LiteLLM
One OpenAI-compatible endpoint, model-aware routing, GPU-aware scheduling, and scaling driven by inference pressure—not CPU alone.
Self-hosting an LLM on Kubernetes is easy to demonstrate. Operating a
blog.shubhamtatvamasi.com26 min read