Eethanwalkerinethanwalkerwrites.hashnode.dev·6d ago · 3 min readA 91% eval pass rate shipped our worst regression. We gate on the delta now.Our CI eval gate failed anything under 90%. A change landed at 91%, passed, and shipped a regression that cost us two days. The absolute threshold was the problem. it cannot tell a stable 91% from a 900
Eethanwalkerinethanwalkerwrites.hashnode.dev·Aug 13 · 8 min read356 of the 528 eval cases we could judge have never failed. Our pass rate could only move 21 points. TL;DR: I broke out per-case results for the incident-harvested part of our eval suite, 611 of our 1,400 cases, and asked which cases have ever discriminated between two shipped versions. Of the 528 wi00
Eethanwalkerinethanwalkerwrites.hashnode.dev·Aug 11 · 10 min readWe back-tested our eval gate against 41 known regressions. It caught 23The gate had been green for eleven weeks. In the same eleven weeks we rolled back two prompt changes, hotfixed a retrieval config on a Saturday, and shipped a summarizer that started dropping the seco00
Eethanwalkerinethanwalkerwrites.hashnode.dev·Aug 6 · 6 min readOur eval gate runs 22 minutes. The queue behind it hit three hours.Nobody complained about the gate. They complained about Tuesday. Our merge queue runs the full eval suite before anything lands: 1,400 cases, about 240 of them scored by an LLM judge, the rest determi00
Eethanwalkerinethanwalkerwrites.hashnode.dev·Aug 4 · 8 min readHalf my eval regressions never touched the prompt fileI keep a tally. Over the last two quarters, a little more than half of the output regressions I traced in one RAG service came from a change that never went near the prompt directory. The eval gate wa00