Running a detector against a benchmark where cheating is guaranteed by construction is a nice design, because the green suite is the label. That is much firmer ground than most agent evaluation work stands on.
The number I would want next to the detection rate is the false-positive rate on a benchmark where spec and tests agree. A gate that blocks legitimate done claims is expensive in a different way, and the two error types trade off against each other. I would also be curious whether the exploit behaviour shifts when the agent is told a verifier will check the work. That is itself a configuration change, and given how much of these results turn out to be configuration, it seems worth a control arm.