Qwen 3.8 Flash on 64GB RAM and 8GB VRAM Custom Fork LLama
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
This is mostly to explain my experience optimizing the model to run on my machine and maybe getting interest from someone to go and write an actual PR against llamacpp (as I'm absolutely not willing to generalize this code ahah). Short Preface (this post is hand written, no AI here). So when Qwen d…
Read the full story at r/LocalLLM ↗
Timeline · 7 reports
- 2026-09-12 11:02 · r/LocalLLaMA
I am impressed and I owe you one, Qwen 3.8 flash next (vision)! - 2026-09-12 07:02 · r/LocalLLM
Qwen flash next beats Fable - 2026-09-12 00:01 · r/LocalLLM
Qwen 3.8-Flash-Next On Mac Mini with 64gb Works - 2026-09-11 22:12 · r/LocalLLM
Minimal total RAM + VRAM to run Qwen 3.8 flash next with decent speed? - 2026-09-11 14:58 · r/LocalLLM
Qwen 3.8 Flash Next MTP tested speed - 2026-09-11 10:40 · r/LocalLLaMA
7900 XTX + 32/64GB RAM for Qwen 3.8 Flash Next? - 2026-09-11 01:28 · r/LocalLLM
Qwen 3.8 Flash on 64GB RAM and 8GB VRAM Custom Fork LLama