Qwen3.8-Flash-Next 125B at 17–26 tok/s on a 16 GB GPU + 32 GB RAM
I’ve been working on running Qwen3.8-Flash-Next on relatively modest hardware: RTX 5060 Ti 16 GB 32 GB DDR5 NVMe SSD The model is ~76 GB, so the experts are split across VRAM, RAM and NVMe. Current real-workload speeds: Code: 26.4 tok/s Agent: 21.3 tok/s Reasoning: 21.8 tok/s 21K context: 17.7 tok/…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-30 18:26 · r/LocalLLM
Qwen3.8-Flash-Next 125B at 17–26 tok/s on a 16 GB GPU + 32 GB RAM