TThrottleinthrottle.hashnode.dev·3d ago · 4 min readThe FP8 trap: my GPU bill dropped 47% because the model was printing "!!!!!!"I had one hour on an AMD MI300X and one question: what does a token actually cost on it? One GPU, vLLM's ROCm build, Qwen2.5 at 7B, 32B and 72B. Thirty-two requests in flight, 256 output tokens max, p00
TThrottleinthrottle.hashnode.dev·Sep 26 · 8 min readI changed nothing and my LLM server got 27% more expensiveI ran the same cost check four times in a row against a server I didn't touch. Same model, same machine, same prompts, same settings. This is what came back, in dollars per million output tokens: R00