IPIvan Portaintodea.hashnode.dev·Sep 13 · 11 min readKubeCon China 2026: Capacity Recovered From The Same HardwareKubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China 2026 in Shanghai ran on two ideas. The first was a workload shift the opening keynote put analyst figures behind: AI compute is n00
IPIvan Portaintodea.hashnode.dev·Aug 27 · 19 min readFrom Ingress to Inference Gateway: Gateway API and the Inference ExtensionTraditional load balancing assumes all backends are the same, so it does not matter which one handles a request. LLM traffic changes this. For example, one vLLM replica might have a warm KV cache for 00
IPIvan Portaintodea.hashnode.dev·Jul 24 · 20 min readLLM Inference Explained: What Actually Happens When You Serve a ModelAs spending on frontier AI services like ChatGPT, Claude, and Gemini climbs, usage caps hit developers, and open-source models gain popularity, most platform teams eventually need to serve an LLM on t00
IPIvan Portaintodea.hashnode.dev·Jul 14 · 6 min readKubeCon India 2026: Sovereign AI On A Cloud-Native StackKubeCon + CloudNativeCon India 2026 in Mumbai ran on one throughline: sovereign AI as an architectural need rather than a slogan, with population-scale platforms like Sarvam and NPCI running on the sa01I
IPIvan Portaintodea.hashnode.dev·Jul 14 · 6 min readLinkerd Explained: The Service Mesh That Stays Out of Your WayLinkerd is an open-source, CNCF-graduated service mesh that adds mTLS, retries, timeouts, and golden-signal metrics to Kubernetes services through a lightweight Rust-based micro-proxy, with no changes00