I benchmarked a few NVFP4 quantizations of Qwen3.8 27B
BUDGET vs MTP-N4_0 vs Q5K-mtp-yarn1m tg128 (token generation) Comparison Median (old → new) Delta p-value BUDGET → MTP-N4_0 31.71 → 30.86 −2.69% 0.000 BUDGET → Q5K-mtp-yarn1m 31.71 → 31.22 −1.53% 0.000 MTP-N4_0 → Q5K-mtp-yarn1m 30.86 → 31.22 +1.19% 0.000 BUDGET is the fastest; Q5K-mtp-yarn1m beats…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-20 02:08 · r/LocalLLM
I benchmarked a few NVFP4 quantizations of Qwen3.8 27B