The guard-vs-feature split in Mike's comment is the right fix, but there's a second source of noise underneath it: the grader itself. If the eval is scored by an LLM judge rather than a fixed assertion, a "refused" verdict can flip to "passed" between runs purely from judge variance, independent of anything the skill under test did. Before locking a must-not-regress gate on a guard case, it's worth measuring the judge's own repeat-run agreement on that exact case, not just the skill's. Otherwise you can build a solid regression gate on a scorer that isn't stable enough to gate on.