"Locker-room mathematics" for parameter counts is a fair description of how model selection often sounds.
The practical argument against size-first is that for most production tasks the model is not the bottleneck. In anything retrieval-shaped, a 7B with the right passage in context beats a 405B with the wrong one, and the gap is not close. Parameter count buys you capability on the hard tail; it does nothing for the far more common failure where the evidence never arrived.
The other cost worth naming is that bigger models make everything slower to iterate on, so you run fewer experiments per week. That compounds against you far faster than a few benchmark points compound for you.