Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
For the next task I'd hold some tests back. Three clean runs on six bugs show the bench works, but visible tests won't separate models much, because a failing test also shows the agent what a passing answer looks like.
Keep giving the agent the six tests, then score against a second file it never sees: the same bugs, checked with different inputs. For an inventory module that could be a negative quantity, or an item that isn't stocked yet. A fix that special-cases the visible inputs passes the first file and fails the second. The gap between those two scores is the number I'd want from a cross-model comparison, and it costs you one more locked hash.
Locking the test file protects the assertions, but the independent rerun should also prove that all six intended tests actually executed against the edited inventory module. A green process exit can still come from empty discovery or importing an installed copy instead of the fresh workspace.
I would record discovered test IDs, executed count and the resolved module path alongside the test hash. Make a zero-test run fail the bench, and retain the six per-test outcomes rather than only Yes in the pilot table. Those checks preserve your small reproducible setup while making its PASS evidence about the intended implementation rather than just the test runner's exit status.
Locking the test file with a SHA-256 and disqualifying any run that touches the hash is the right instinct, because the moment the agent can edit the scorer, PASS means nothing. The n=3 honesty is good to see too, most people would have shipped a ranking. I ran into the other half of this, a suite that stayed green while the agent still broke in production, and wrote up the CI/CD I moved to: kartiknvjk.hashnode.dev/my-agent-passed-every-che… . Are you adding held-out hidden tests on top of the locked file?