nvfp4 vs k8v4 KV cache on Qwen3.8-27B @ 224K ctx
This story is from 2026-09-13. It is preserved in the archive; the latest stories are on the live feed.
Ninfer recently added NVFP4 KV cache support, so I decided to A/B test to see if I can switch to nvfp4 kv for more ctx, was using k8v4. Model: Qwen3.8-27B nvfp4, 224K context, temp 1.0, thinking on. A/B diff: --kv-dtype k8v4 vs --kv-dtype nvfp4. 1) Needle pickup with fake needles : a 224K document…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-13 16:55 · r/LocalLLM
nvfp4 vs k8v4 KV cache on Qwen3.8-27B @ 224K ctx