Building the evaluation set before comparing models is the step that separates this from leaderboard shopping, and it is also the step everyone tries to skip because it is the only genuinely tedious part. One thing I would add to the multi-objective section: run the comparison at the same prompt effort per model rather than the same prompt. A prompt tuned for one family often underperforms on another for reasons that have nothing to do with capability, and the resulting table looks like a capability gap. The other practical constraint that rarely makes these frameworks is switching cost - a model that is five percent better but changes your output format, your latency profile and your failure modes is not five percent better once you price the migration and the regression risk.