Worth pushing on the 405B-vs-Pythia example a bit further: that LLaMA was 4-bit quantized while the Pythia models weren't, and outlier features (the ones LLM.int8() had to special-case) get more pronounced as models scale up. Naive quantization degrades a big model's effective capability more than it does a small one's, so there's a second confound sitting right on top of the ones already listed — a fair size comparison needs matched precision, not just matched task. Ran into this exact trap comparing an MXFP8 128GB local model against smaller full-precision ones; had to control for the quantization scheme before any size conclusion was trustworthy.