Rule 10 is the one I'd want to see more of elsewhere, naming specific cheat shapes (vacuous success, playing to the grader, one-direction equivalence claims) turns "be suspicious" into something you can actually audit mechanically instead of relying on a reviewer remembering to be paranoid that day. Most AI-verification setups stop at "have a human check it," which doesn't scale past cycle 10, let alone 104.
The G-105 refutation being reframed as the first machine-checked evidence for a real claim is the best illustration of Rule 3 actually paying off, that only works because the failure-policy format was fixed before the refutation happened, not improvised afterward when the null result showed up.
Rule 12 pairing with Rule 13 is the quieter but more important design choice though, banning full-project verification inside the loop only works because a human explicitly decided where the cost tradeoff sits, rather than the loop discovering on its own that thorough checking is expensive and cutting corners.