Qwen flash next on 12+16gb vram, and 32gb ram viable?
I have 4070 12gb, and v100 16gb, and 32gb 3200mhz ram, with the basic llama, with flash attention 98k ctx on q4, with layer split NOT tensor parallelism, as the v100 is sxm2 and has a adapter with pcie 3.0 x16 and I have pcie 4.0 x4, so tensor parallelism slows this down by like .7 tps, Im on ubunt…
Read the full story at r/LocalLLM ↗
Timeline · 5 reports
- 2026-10-01 05:27 · r/LocalLLaMA
Qwen Flash Next MTP work restarted - 2026-09-30 17:15 · r/LocalLLaMA
How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks? - 2026-09-30 14:29 · r/LocalLLM
Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really) - 2026-09-30 09:41 · r/LocalLLM
Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo - 2026-09-29 10:21 · r/LocalLLM
Qwen flash next on 12+16gb vram, and 32gb ram viable?