Best current Qwen Flash Next Q4-ish? + worth using?
Im running a 5090 and 64gb of ram, so im limited on what I can run. I have currently been able to fit the following - Atomic Q4_k_m 4.27bpw @ 31 layers offload Swift IQ4_xs @ 32 layers offload. Im about to try the Unsloth IQ4_xs as well. I could get a "bigger" (non IQ) quant for atomic because its…
Read the full story at r/LocalLLaMA ↗
Timeline · 6 reports
- 2026-09-28 07:49 · r/LocalLLaMA
If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context. - 2026-09-27 20:18 · r/LocalLLM
Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system - 2026-09-26 21:11 · r/huggingface
Abbiamo inserito 100 informazioni nella tabella engrammatica di Qwen 3.8 Flash Next e abbiamo creato un sito web per illustrarle. - 2026-09-26 19:54 · r/LocalLLaMA
Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much? - 2026-09-26 13:02 · r/LocalLLM
We have implanted 100 facts into the engram table of Qwen 3.8 Flash Next, and we have now created a website to explain it. - 2026-09-26 08:52 · r/LocalLLaMA
Best current Qwen Flash Next Q4-ish? + worth using?