The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell
If you run vLLM with `--kv-cache-dtype fp8` on a DeepSeek-family (MLA) model and your GPU is a GB10, an RTX PRO 6000, or any workstation or consumer Blackwell card, there is a decent chance the engine