Partly, here is how i have it: the question wording and the thresholds live in one place, and changes to them get scored against labeled sets. One synthetic set with planted stale sentences and opposing claims the excerpt does not support. One real set built from my own past corrections, where the original sentence counts as stale and the corrected form as fine. The same documents go through two framings of the question and I read the union, because they flag different sentences. Before I trust a model at all it answers three sharp controls: obviously true, obviously false, unrelated. And real tasks get planted counterexamples. The stale index entry was one of those plants. It was the adversarial pass, and it failed, which is why indexes stayed manual. "Nothing found" is not green in my setup, it is just the answer distribution. What I don't have is an adversarial pass on the fact sheet itself. It is hand-written from dated project records, and its errors surfaced exactly the way you describe: false alarms that an agent traced back to a fact written too broadly. The fact sheet gets reviewed like any other document, by an agent and by me, not by the judge. On "can't tell unsupported from contradicted": the most recent version of that I hit with a local model was not wording, it was evidence width. Same claim, same model. Against a one-line status it answers "contradicted" at full confidence. Against the 600-character head of the same document it answers "not established by this evidence", also at full confidence, and that lands in the review band, so anyone reading only flagged items sees nothing. So the rule is now one assertion per question against a focused excerpt that still keeps the scope, the dates and the exceptions. That is a heuristic, not a guarantee. And when the evidence is silent on something the claim asserts, I write "not measured" into the evidence instead of hoping the judge infers it. The judge has no permissions, and the other agent reviews the checking tools, not just the documents.