The "not a ranking of intelligence" framing is worth stating explicitly, most setups drift the other way without meaning to. Our own pipeline runs an adversarial review pass before anything gets published, and the temptation is always to treat that reviewer as the senior one instead of a second set of eyes with a narrower question. The permission point is the actual teeth of it though: your judge has no publish/approve rights regardless of how confident it sounds. That's the real guardrail, not its accuracy.
The "not a ranking of intelligence" framing is worth stating explicitly, most setups drift the other way without meaning to. Our own pipeline runs an adversarial review pass before anything gets published, and the temptation is always to treat that reviewer as the senior one instead of a second set of eyes with a narrower question. The permission point is the actual teeth of it though: your judge has no publish/approve rights regardless of how confident it sounds. That's the real guardrail, not its accuracy.
Clean way to name what a lot of us do ad hoc. The part I'd push on: judge-has-no-write-permission is the right default. But the harder failure is when the judge's narrow question is wrong itself. Bad fact sheet. A question that can't tell "unsupported" from "contradicted." Nothing downstream catches that except a human noticing the scan missed something obvious. Do you run any adversarial pass on the fact sheet and question wording themselves, or does that only get caught when a defect slips through like your stale index entry? We've hit exactly that in our own review gate: a reviewer confidently answering a badly-framed question, which is worse than a model lying about a good one.