MXFP8 vs Q8: 10x the weight error, 1% the perplexity
MXFP8 vs Q8: 10x the weight error, 1% the perplexity
Key takeaways
At the same 8.25 bits per weight, MXFP8 reconstructed real Qwen3-Coder-30B weights with 6.85% mean layer-output error against 0.67%
leanzero.hashnode.dev19 min read