The most dangerous failure mode here is the “valid but meaningless” result. A parser accepting a property or a compiler accepting code only proves that the structure is legal; it says nothing about whether the system actually changed in the intended way.
I especially like the distinction between checking the explanation and checking the observable outcome. Version-pinning, reducing to a minimal reproduction, diffing the result, and verifying against the primary source create a much stronger validation loop than asking another model whether the first answer looks correct.
The silent-failure numbers are a useful reminder too: when AI-generated technical advice can fail without producing an error, testing isn't the final step after reasoning it has to be part of the reasoning loop itself.