AKAyesha Khaninayeshakoder.hashnode.dev·8h ago · 7 min readBelievable Is Not CorrectOur owner wants an AI agent in the app. It should understand the regional language, take voice notes, cope with different dialects, and do real work inside the product. It's an exciting request, and I00
OZOliver Zehentleitnerinblog.technopathy.club·4d ago · 9 min readThe Model Changed. My Skill Didn't. The Score Still Dropped.My rule for evaluating an agent skill is deliberately asymmetric: Test the agent on the weakest model you intend to support. Choose the judge by measuring which model grades that task reliably. Those 00
SSanathinaiqa.hashnode.dev·Sep 29 · 3 min readAdding a Tool to Your AI Agent? Here's What Actually Needs RetestingTwo changes account for most of the "wait, when did that break" moments teams have with AI agents: adding a tool or widening a permission, and updating the model, system prompt, or orchestration layer01M
LSLakshaya Sharmainblog.langersword.in·Sep 11 · 12 min readyour agent passed the eval. it still failed the job.Most AI projects add evals in roughly the same way. Someone picks a benchmark, collects a set of prompts, runs a model, and turns the results into a percentage. The percentage becomes a line on a dash00
OZOliver Zehentleitnerinblog.technopathy.club·Aug 24 · 15 min readSame Skill, Six Agents, Nine Models: What a Real Eval Matrix Taught MeI tested Keep the Why, my open-source agent skill for preserving the reasoning behind a codebase, through six coding agents, with a matrix spanning nine hosted models plus a local Ollama run. The mode20
MSManu Shuklainecorpit.hashnode.dev·Aug 20 · 14 min readOpenAI shuts Agent Builder, Evals and reusable prompts on 30 November 2026OpenAI shuts Agent Builder, Evals and reusable prompts on 30 November 2026 Summary. OpenAI is retiring three developer-platform surfaces on the same day. Agent Builder, the Evals platform and reusable00
SGSunny Guptainicodestartups.com·Jul 28 · 7 min readSeven Models in Seven Days: How to Pick Without WhiplashBetween July 17 and July 23, the AI world shipped seven notable models in seven days. Kimi K3 with a trillion parameters. Three separate Qwen drops. Google's Gemini 3.6 Flash family. poolside's open-w21N
KDKrishna Dev Paleminkrishnadevpalem.hashnode.dev·Jul 20 · 13 min read Should an AI Decide What Gets Erased? I Measured It.I built a system that adjudicates data-erasure requests with no model in the decision path, then built an evaluation harness to measure what a capable model would have done in its place. The model alm00
JLJeremy Longshoreinjeremylongshore.hashnode.dev·Jul 10 · 18 min readNoise-Robust LLM-Judge Evals: Don't Sign a Coin FlipAn un-seeded LLM judge is nondeterministic even at temperature 0. So a single-call binary verdict is a coin flip. And when you sign that verdict into a public transparency log, you have cryptographica00
ACAlexander Codesinalexandercodes.hashnode.dev·Jun 29 · 10 min readHow to Evaluate AI-Generated SQL on DatabricksAn AI assistant can produce a query that runs and returns a plausible result. But how do we know if it's correct? Suppose a user asks: Which enterprise customers had revenue growth last quarter? An 00