STShubham Tatvamasiinblog.shubhamtatvamasi.com·6h ago · 26 min readRunning LLMs on Kubernetes in Production: KServe, vLLM, llm-d, Envoy AI Gateway, and LiteLLMOne OpenAI-compatible endpoint, model-aware routing, GPU-aware scheduling, and scaling driven by inference pressure—not CPU alone. Self-hosting an LLM on Kubernetes is easy to demonstrate. Operating a00
STShubham Tatvamasiinblog.shubhamtatvamasi.com·Jun 14, 2025 · 2 min readBuild for 10,000 YearsMost people working in tech today don’t think about legacy. They focus on finishing the sprint, shipping the next release, or just getting through the week. Their contribution often stops at what’s assigned, and rarely do they pause to understand the...00