Comparing Two Eval Runs by Their Average Pass Rate Is the Wrong Test
TL;DR. You run version A and version B against the same 500-item eval set. A passes 71.4 percent, B passes 74.0 percent, and you conclude B is better. That reasoning throws away the one fact that matt
llmasajudge.hashnode.dev18 min read