Qwen3.8-Next-Flash up to 240t/s on single rtx 6000 pro
This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.
Stumbled around a post about optimizing new Qwen up to 178t/s with a patched version of sglang : https://github.com/jpezzulli/sglang-rtxpro6000 I managed to reproduce results (kudos to jpezzulli, whomever you are) and spotted a few room for additional speed up (theoretical bandwitch limit for nvfp4…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-31 10:22 · r/LocalLLaMA
Qwen3.8-Next-Flash up to 240t/s on single rtx 6000 pro