Good that you're skeptical of the 98%-retention number. Those figures are almost always benchmark averages weighted toward knowledge/reasoning tests like MMLU, which measure something pretty different from a from-scratch generate-and-iterate task like the Three.js build. A model can score close to the original on one axis and still fail completely on the other, which is closer to what you actually hit over those three hours than a straight 2% quality gap would predict.
Yeah, I was pretty excited in the beginning, but after spending 3-4 hours on that Three.js generation task, I stepped back a little. Still, impressive.
Mihai_LeanZero
Atlassian consultant and Forge developer at LeanZero. We do Data Center to Cloud migrations and build Forge apps, and I write up what I learn at leanzero.net.
Good that you're skeptical of the 98%-retention number. Those figures are almost always benchmark averages weighted toward knowledge/reasoning tests like MMLU, which measure something pretty different from a from-scratch generate-and-iterate task like the Three.js build. A model can score close to the original on one axis and still fail completely on the other, which is closer to what you actually hit over those three hours than a straight 2% quality gap would predict.