With one planned repeat, the invalid-date flip is a single sample, so the evidence can't yet tell a regression from nondeterminism. I'd rerun just the changed assertions several times each before deciding what the flip means.
The bigger problem is the equal-weight mean. Refusing an invalid date and getting the report total right aren't the same kind of assertion: one is a feature, the other is a guard. I'd tag guard assertions as must-not-regress and have them gate on their own, outside the average. Then a candidate that trades a guard for a feature fails on its face, and the 75% becomes a summary of everything else rather than the release decision.
Your point about checking the output behind the grade matters most for the guards. A grader that reports "refused" when the Skill quietly coerced the bad date into a valid one is the failure I'd look for first.