Qwen3.8 27B runs at ~480 tok/s on a single 5090
Been doing some testing, swapped NInfer's official Qwen3.8-27B quant (part NVFP4, part FP8) for QUASAR's QAT full-NVFP4 checkpoint, with DFlash2 embedded. Running on the same hardware and engine flags: metric official this JSON output 400 tok/s 484 tok/s prose 176 tok/s 204 tok/s prefill 11.4k tok/…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-01 16:30 · r/LocalLLM
Qwen3.8 27B runs at ~480 tok/s on a single 5090