Mminhleducinminhleduc.hashnode.dev·Aug 16 · 8 min readWhy Production Agents Are Ditching Pure RAG for Context CachingThree months ago, our agentic workflow hit a cost wall. We were running a multi-step research agent that performed iterative document analysis. Every user request triggered between 12 and 18 agent ste00
Mminhleducinminhleduc.hashnode.dev·Aug 8 · 15 min readKimi K3 has 2.8T parameters. Moving them was the hard part.Every article about Kimi K3 leads with the same number. 2.8 trillion parameters, the largest open-weight model anyone has shipped. It's a real number and it's the least interesting thing in the techni00
Mminhleducinminhleduc.hashnode.dev·Aug 8 · 12 min readHow Kimi K3 reads a million tokensStandard attention has one genuinely magical property: any token can look directly at any earlier token, at full fidelity, no matter how far back it sits. Nothing is summarised. Nothing is forgotten. 00
Mminhleducinminhleduc.hashnode.dev·Aug 5 · 12 min read5 Tiny Language Models for Tool Calling (Part 3)The benchmark nobody quotes when they tell you small models are the future. NVIDIA Research published a position paper called "Small Language Models are the Future of Agentic AI". I believe the argume00
Mminhleducinminhleduc.hashnode.dev·Aug 5 · 10 min read5 Tiny Language Models You Need to Know in 2026 (Part 2)I wrote Part 1 in October 2024. The subtitle was "Tiny is a new large!" and at the time that felt like a slightly cheeky thing to say. It doesn't anymore. Google now ships a model that runs in 1.1 GB 00