The design choice hiding in your setup deserves top billing: the reviewer had access to tests the author's environment could not run, and you call that "useful, not inconvenient." That asymmetry is what makes "do not use a second AI as another voter. Use it as an independent source of friction" actually work. Two agents with identical checkouts, tools and context converge on identical blind spots, and the vote is theater; the reviewer that caught your ten findings caught them by running things the author physically could not. Which suggests a stronger version of your protocol: engineer the asymmetry deliberately, different tool access, different test tiers, maybe a different model family, so agreement carries information. The finding that the hardest round ended with less code, deleting the semantic detector instead of patching it a fifth time, is also the most human-senior-engineer behavior I have seen attributed to a review loop.
The design choice hiding in your setup deserves top billing: the reviewer had access to tests the author's environment could not run, and you call that "useful, not inconvenient." That asymmetry is what makes "do not use a second AI as another voter. Use it as an independent source of friction" actually work. Two agents with identical checkouts, tools and context converge on identical blind spots, and the vote is theater; the reviewer that caught your ten findings caught them by running things the author physically could not. Which suggests a stronger version of your protocol: engineer the asymmetry deliberately, different tool access, different test tiers, maybe a different model family, so agreement carries information. The finding that the hardest round ended with less code, deleting the semantic detector instead of patching it a fifth time, is also the most human-senior-engineer behavior I have seen attributed to a review loop.