I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
I tested two local LLMs in two different quantization formats on the same laptop, on the same prompt, three trials each. The result is the opposite of what the marketing says: the supposedly-faster new format lost by 1.8x. Q4_K_M (the older integer-based format) hit 4.7 tokens/second and finished a…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-05 01:47 · DEV Community — AI
I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost