Building an NVFP4 KV Cache for a Hybrid Qwen Model
Packing K/V into 576 bytes per token, fixing vLLM's hybrid cache planner, and measuring the result on an RTX PRO 6000 Blackwell. I spent a fair amount of time getting NVFP4 KV storage working in my Qwen3.8-Flash-Next serving stack. The work covered the quantizer, a packed-cache writer, a sparse att…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-27 14:33 · DEV Community — Machine Learning
Building an NVFP4 KV Cache for a Hybrid Qwen Model