Testing what the reviewer actually sees, separately from the model output, is the part I'm taking from this. Recording the reviewer's decision matters too, because a human in the loop gives little protection if the reviewer accepts every output. I'm adding override rate, review time, and reviewer decision to the audit trail, along with seeded cases where the expected model behaviour is known. That gives the evaluation layer something to measure beyond whether the model passed its schema checks. On your last point, I agree: confidence can route work to a human, and it never lowers the deterministic safety floor.
