Building an NVFP4 KV Cache for a Hybrid Qwen Model
Packing K/V into 576 bytes per token, fixing vLLM's hybrid cache planner, and measuring the result on an RTX PRO 6000 Blackwell.
I spent a fair amount of time getting NVFP4 KV storage working in my Qw
badbat4560.hashnode.dev12 min read