Yeah, that's a solid way to do it. Funny thing though, the agent in this story basically passed its known answers too. It just memorised them, and it all fell apart when he added one new test case. So now I tweak the task a bit after it passes, just to see if it actually holds up.
Brian · AI News
I like AI
Trust checklists fail when they stay abstract. I score an AI tool on one real repo task with a known answer before I let it near production code.