Thanks, Mike and Mihai. I agree that a recorded flip is not yet evidence of a repeatable regression.
The article uses authored synthetic records to show how equal averages can hide different outcomes. It doesn't establish the cause or frequency of a failure.
For a real evaluation, I'd separate three checks: inspect the raw outputs; measure behavior across planned repeats; and, when using an LLM judge, rescore fixed outputs to check the judge's agreement. I'd keep any targeted diagnostic reruns separate from the original comparison.
The guard/feature distinction should be decided before looking at the candidate's results. A critical guard shouldn't disappear inside an average, but its scoring rule needs to be dependable too—especially for the difference between rejecting an invalid date and silently coercing it.
As the checker's maintainer, I should be clear about its boundary: it reviews supplied evidence; it doesn't verify the grader, establish statistical reliability or choose the release gates. Those remain separate checks.