Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
Thanks, Mike and Mihai. I agree that a recorded flip is not yet evidence of a repeatable regression.
The article uses authored synthetic records to show how equal averages can hide different outcomes. It doesn't establish the cause or frequency of a failure.
For a real evaluation, I'd separate three checks: inspect the raw outputs; measure behavior across planned repeats; and, when using an LLM judge, rescore fixed outputs to check the judge's agreement. I'd keep any targeted diagnostic reruns separate from the original comparison.
The guard/feature distinction should be decided before looking at the candidate's results. A critical guard shouldn't disappear inside an average, but its scoring rule needs to be dependable too—especially for the difference between rejecting an invalid date and silently coercing it.
As the checker's maintainer, I should be clear about its boundary: it reviews supplied evidence; it doesn't verify the grader, establish statistical reliability or choose the release gates. Those remain separate checks.
The guard-vs-feature split in Mike's comment is the right fix, but there's a second source of noise underneath it: the grader itself. If the eval is scored by an LLM judge rather than a fixed assertion, a "refused" verdict can flip to "passed" between runs purely from judge variance, independent of anything the skill under test did. Before locking a must-not-regress gate on a guard case, it's worth measuring the judge's own repeat-run agreement on that exact case, not just the skill's. Otherwise you can build a solid regression gate on a scorer that isn't stable enough to gate on.
With one planned repeat, the invalid-date flip is a single sample, so the evidence can't yet tell a regression from nondeterminism. I'd rerun just the changed assertions several times each before deciding what the flip means.
The bigger problem is the equal-weight mean. Refusing an invalid date and getting the report total right aren't the same kind of assertion: one is a feature, the other is a guard. I'd tag guard assertions as must-not-regress and have them gate on their own, outside the average. Then a candidate that trades a guard for a feature fails on its face, and the 75% becomes a summary of everything else rather than the release decision.
Your point about checking the output behind the grade matters most for the guards. A grader that reports "refused" when the Skill quietly coerced the bad date into a valid one is the failure I'd look for first.
This is the failure mode aggregate scores are built to hide, both runs read 75% while one assertion flipped FAIL to PASS and another went the other way, net zero on the dashboard, real regression underneath. I got paged at 3 AM by exactly this, green CI, regression still live, and wrote up the per-assertion diffing I switched to: kartiknvjk.hashnode.dev/my-ci-evals-were-green-a-…. Do you alert on per-assertion deltas, or is the single score still the gate?