For the next task I'd hold some tests back. Three clean runs on six bugs show the bench works, but visible tests won't separate models much, because a failing test also shows the agent what a passing answer looks like.
Keep giving the agent the six tests, then score against a second file it never sees: the same bugs, checked with different inputs. For an inventory module that could be a negative quantity, or an item that isn't stocked yet. A fix that special-cases the visible inputs passes the first file and fails the second. The gap between those two scores is the number I'd want from a cross-model comparison, and it costs you one more locked hash.