Scoring retrieval and answer faithfulness separately is the decision that makes this harness trustworthy, since a faithful answer sitting on lucky retrieval will pass a naive check and then fail the moment the corpus shifts. The result that stood out to me is the refusal rate: 4 out of 8 unanswerable questions refused means the guardrail prompt is closer to a coin flip than a guarantee, which no amount of manual spot-checking would have surfaced. When your judge flags a hallucinated claim, do you feed those failing cases back as fixed regression items so a prompt tweak that fixes one refusal cannot silently break another?
Scoring retrieval and answer faithfulness separately is the decision that makes this harness trustworthy, since a faithful answer sitting on lucky retrieval will pass a naive check and then fail the moment the corpus shifts. The result that stood out to me is the refusal rate: 4 out of 8 unanswerable questions refused means the guardrail prompt is closer to a coin flip than a guarantee, which no amount of manual spot-checking would have surfaced. When your judge flags a hallucinated claim, do you feed those failing cases back as fixed regression items so a prompt tweak that fixes one refusal cannot silently break another?