The point about who spends the time is sharp. Familiarity acts like a hidden multiplier on the fixed budget, so the same hour buys different amounts of progress depending on which model the engineer knows best. The article's framework assumes the evaluation data does the heavy lifting, but your observation shows the human in the loop still tilts the results. A simple fix is to have each engineer write a quick log of what they tried and why. That log often reveals whether the model is good or the engineer is just good at steering it.