AJAlen Joyinpragmaticstack.in·7m ago · 16 min readThe Better Model Broke Production: Rolling Out a New Model Safely🗓️ Last updated: July 2026 The new model revision cleared the eval suite on a Monday. The numbers were unambiguous: faithfulness up three points, hallucination rate down, latency within budget. Shad00
AJAlen Joyinpragmaticstack.in·5d ago · 15 min readThe Behaviour Was the Bug: Observability for Systems Where Stack Traces Don't Help🗓️ Last updated: July 2026 A support ticket arrives on a Monday morning. A user asked the assistant a billing question. The answer was wrong, confident, and specific enough to do damage. The user fo43MM
AJAlen Joyinpragmaticstack.in·Jul 14 · 15 min readThe Diff Was Empty: Treating Prompts as Code and Models as Dependencies🗓️ Last updated: July 2026 The release diff is empty. No commits since Tuesday. No infrastructure changes. The dashboards from a week ago and the dashboards from this morning track different distrib00
AJAlen Joyinpragmaticstack.in·Jul 9 · 15 min readYour Eval Lied: The Architecture of a Real LLM Evaluation Program🗓️ Last updated: July 2026 The team spent six weeks on a model upgrade. Every internal eval is green. The new model beats the old one on the golden set by four points. It beats it on the regression 10
AJAlen Joyinpragmaticstack.in·Jul 7 · 21 min readThe Agent Reasons, the Engine Remembers: Designing an Agentic Workflow Engine🗓️ Last updated: July 2026 Picture a workflow that began an hour ago with a customer email. Step 1 classified intent. Step 2 retrieved order history. Step 3, an LLM agent, read both, decided the cus31J