Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
Strong write-up. The anchor-set point is the part I’d underline: once the judge is also a model, the harness needs its own calibration surface, not just a better judging prompt.
One enhancement I’d consider is making the eval receipt explicit for every run: judge version, candidate model, rubric version, anchor-set agreement, order-shuffle result, drift signal, and which failures were deterministic versus judgment-based. That makes the dashboard more inspectable when a score changes.
Affiliation note: I’m with nxus.SYSTEMS. This overlaps with nxusKit SDK CE examples around model-research harnesses, structured output, deterministic checks, Bayesian confidence, and retry/fallback patterns. CE is always free.
The easiest way to find the examples is to search for: nxusKit SDK examples
The section on judge model drift is the most important thing in this post. I have seen teams set a CI gate on judge scores, then upgrade the judge model a month later and watch every number shift without the actual agent changing. The fix is to pin the judge model and re-baseline after every upgrade. I wrote about this in my series on evaluating OpenAI Agents SDK apps hashnode.com/edit/cmshzhzgo00000bj99poe7yi9 How often do you refresh your judge calibration set?