Good that you're skeptical of the 98%-retention number. Those figures are almost always benchmark averages weighted toward knowledge/reasoning tests like MMLU, which measure something pretty different from a from-scratch generate-and-iterate task like the Three.js build. A model can score close to the original on one axis and still fail completely on the other, which is closer to what you actually hit over those three hours than a straight 2% quality gap would predict.