Capping human iteration time per candidate is a better control than a shared prompt, and I had not framed it that way before. It measures what a team will really spend rather than an idealised equal starting point.
The thing I would guard is who spends the time. Familiarity with a family buys a lot inside a fixed budget, so the same hour goes further on the model the engineer already knows. With two people, swapping which candidate each one tunes is a cheap way to stop that quietly becoming the result. Recording the trajectory rather than only the final score helps too: a candidate still improving when the budget expires is telling you something quite different from one that plateaued in twenty minutes.
Building the evaluation set before comparing models is the step that separates this from leaderboard shopping, and it is also the step everyone tries to skip because it is the only genuinely tedious part. One thing I would add to the multi-objective section: run the comparison at the same prompt effort per model rather than the same prompt. A prompt tuned for one family often underperforms on another for reasons that have nothing to do with capability, and the resulting table looks like a capability gap. The other practical constraint that rarely makes these frameworks is switching cost - a model that is five percent better but changes your output format, your latency profile and your failure modes is not five percent better once you price the migration and the regression risk.