yeah, the empty folder one is brutal because everything upstream looks fine. that's the whole problem with checking the summary.
for non-code work i score the end state, not what the agent says it did. "updated the docs" really means a file should be different, so the check goes and reads the file: did the target section change in the diff, does it actually contain the thing, any leftover TODO, does it still render. none of that needs a model.
the part that bites is writing the task so it has an end state you can look at. "update the docs" doesn't. "the auth section should describe the new token flow" does, you just go read that section. when i can't phrase it that way i take it as the same signal you did when you killed the handoff: no observable state proving it happened means the step is defined wrong, not just unchecked.
the empty folder version of this is the one that got us. building grunz (a coding agent), we had a setup where one model planned and a second model did the actual writing, and sometimes the second one just quietly didn't write anything. the run still wrapped up looking finished and you'd open the project folder and it was empty. your line about the summary being what the agent believes, not what happened, is exactly it. we ended up killing the handoff instead of adding a check, which in hindsight was fixing it the lazy way. what does your scoring function look like for work that isn't code, like "updated the docs"?