DFlash Changes What Tokens per Second Means
I spent a night trying to fit a dense 30B model, 256K context, vision,
and speculative decoding onto one 24 GB GPU. The fastest quant lost. The
quant with the lowest perplexity lost too. What won was
piszczek.hashnode.dev15 min read