218 tok/s Qwen3.8-27B on a single PCIe V100 32GB
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
I've been seeing how far I can push my V100 with a purpose-built inference engine targeting just Qwen and the v100. Current result with Qwen3.8-27B and using a software NVFP4 that works on this hardware: 218 tok/s best-case MTP decode (99.2% acceptance) ~200 tok/s without cheating or context-copy s…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-11 17:41 · r/LocalLLM
218 tok/s Qwen3.8-27B on a single PCIe V100 32GB