Claude stops when the work looks done. It judges done from the same conversation that produced the code. Nothing in that context is independent of the claim. Move the check out of the chat, a hook or a script that runs on a fresh checkout and can't see the reasoning, and "looks done" is no longer something it can conclude. Same agent, same model, just judged by something that wasn't there when it decided.