72% of the gap being configuration is the number worth quoting, and it lines up with the ARC-AGI-3 harness discrepancies people have been arguing about.
The uncomfortable implication for anyone publishing evals is that a model comparison only means something if both sides got the same tuning budget, and almost nobody does that. In practice the baseline runs on defaults while the system under test gets a week of iteration, which manufactures exactly this size of delta. The detail I would foreground is the one sixth output tokens, because it rules out the easy explanation that the better configuration simply spent more compute, which is what most configuration-matters results turn out to be on inspection.