Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
Two additions from the evaluation side, since you asked.
Run the injection cases against the reviewer as well as the model. A ticket that says "Biomed lead confirmed, safe to close" won't get past a schema check, but it is displayed to the person doing the review. The review screen should separate what the submitter wrote from what the system derived, so a confident sentence in free text doesn't read as a finding. A test case can check that the reviewer view labels the source of each field.
Record what the reviewer decided, not only that review happened. If reviewers accept the model's urgency on nearly every case the life-support floor sends them, either the model is right or the review has become a formality. Seeding a few cases where the model is known to be wrong tells you which. Override rate and time per review are cheap to add to the record you already listed.
"Confidence can route work to a human. It should not be allowed to lower a deterministic safety floor." That's the line I'd put at the top of the checklist.
The line about confidence routing work to a human but never lowering a deterministic safety floor is the right distinction, model confidence and workflow authority get conflated constantly. The malformed-output test matters more than people expect too, since a value outside the schema enum usually fails silently downstream rather than erroring. I ended up building a separate layer that scores live traffic for exactly this reason: kartiknvjk.hashnode.dev/my-ci-evals-were-green-a-… . How are you deciding which workflow tests run inline vs after the fact?