The claim that a judge can hit test-retest stability above 0.95 and still be severely biased is the trap I've watched teams fall into, consistency gets mistaken for correctness. A deterministic floor makes sense for the checks that can be made objective, though plenty of quality dimensions never reduce to a rule you can write down. Where do you draw the line between what earns a hard deterministic check and what has to stay a judged call?