PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill
to all dual 3090s users that want to run flash next FAST!: https://huggingface.co/albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE amazing job by this guy. latest update gives me 80tps in real world scenarios and 2k+ prefill. I personally reduced the context to 220k and increased the hot experts to 88 for…
Read the full story at r/LocalLLM ↗
Timeline · 6 reports
- 2026-09-27 20:18 · r/LocalLLM
Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system - 2026-09-26 21:11 · r/huggingface
Abbiamo inserito 100 informazioni nella tabella engrammatica di Qwen 3.8 Flash Next e abbiamo creato un sito web per illustrarle. - 2026-09-26 19:54 · r/LocalLLaMA
Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much? - 2026-09-26 13:02 · r/LocalLLM
We have implanted 100 facts into the engram table of Qwen 3.8 Flash Next, and we have now created a website to explain it. - 2026-09-26 08:52 · r/LocalLLaMA
Best current Qwen Flash Next Q4-ish? + worth using? - 2026-09-26 03:42 · r/LocalLLM
PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill