Right, and that is the sharper version: matching compute does not make a comparison fair when the asymmetry is human iteration. Compute is at least measurable, which is probably why everyone reaches for it.
Tuning effort has no unit. The closest practical proxy I have seen is capping wall-clock iteration time per candidate, which at least mirrors what a team would actually spend. Its own bias is that the same hour buys more on the model family the engineer already knows, so with two people it is worth swapping who tunes which candidate.
None of that makes the number clean, but reporting the budget next to the score at least makes the residual bias visible instead of invisible.
72% of the gap being configuration is the number worth quoting, and it lines up with the ARC-AGI-3 harness discrepancies people have been arguing about.
The uncomfortable implication for anyone publishing evals is that a model comparison only means something if both sides got the same tuning budget, and almost nobody does that. In practice the baseline runs on defaults while the system under test gets a week of iteration, which manufactures exactly this size of delta. The detail I would foreground is the one sixth output tokens, because it rules out the easy explanation that the better configuration simply spent more compute, which is what most configuration-matters results turn out to be on inspection.