Did FP8 make the model dumber? A per-prompt regression check for quantized serving
This story is from 2026-08-25. It is preserved in the archive; the latest stories are on the live feed.
FP8 gave us a clean 1.5x on Qwen3-8B serving throughput on an RTX PRO 6000 Blackwell (1,725 to 2,597 tok/s at concurrency 32, vLLM). The uncomfortable question is always the same: did the model get dumber. This post is the exact check we ran before recommending the switch, with numbers, so you can…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-08-25 18:12 · DEV Community — Machine Learning
Did FP8 make the model dumber? A per-prompt regression check for quantized serving