The editor's note is the most credible thing on this page. Publishing that the first version linked dead sources and described a vendor-run comparison as a third-party leaderboard result costs something, and it is exactly what makes the rest of the analysis readable.
One thing I would sharpen. A vendor-run 75.4% is not a weaker version of a leaderboard 75.4%. They are different measurements. A public leaderboard fixes the harness, the scaffold and the retry policy; a vendor running its own evaluation chooses all three, and on agentic benchmarks the harness moves scores by more than most model upgrades do. So the two numbers are not comparable even in principle, rather than comparable but pending verification. That distinction matters because the second framing implies the listing will eventually settle the question, and it will not: a later listed result would be a different measurement again.
Which is what makes your four locked variables the real contribution here. If a team could only demand one before migrating, I would argue for the agent harness, because that is where the unstated difference between two runs of the same model almost always lives.