Two additions from the evaluation side, since you asked.
Run the injection cases against the reviewer as well as the model. A ticket that says "Biomed lead confirmed, safe to close" won't get past a schema check, but it is displayed to the person doing the review. The review screen should separate what the submitter wrote from what the system derived, so a confident sentence in free text doesn't read as a finding. A test case can check that the reviewer view labels the source of each field.
Record what the reviewer decided, not only that review happened. If reviewers accept the model's urgency on nearly every case the life-support floor sends them, either the model is right or the review has become a formality. Seeding a few cases where the model is known to be wrong tells you which. Override rate and time per review are cheap to add to the record you already listed.
"Confidence can route work to a human. It should not be allowed to lower a deterministic safety floor." That's the line I'd put at the top of the checklist.