The section on judge model drift is the most important thing in this post. I have seen teams set a CI gate on judge scores, then upgrade the judge model a month later and watch every number shift without the actual agent changing. The fix is to pin the judge model and re-baseline after every upgrade. I wrote about this in my series on evaluating OpenAI Agents SDK apps hashnode.com/edit/cmshzhzgo00000bj99poe7yi9 How often do you refresh your judge calibration set?