Qwen3.8-Flash-Next (95.5 GiB) on a 64GB Mac at ~27 tok/s, checkpoint + fork
I've been running Qwen3.8-Flash-Next as my main local coding model from past few weeks. The file is 95.5 GiB and my Mac has 64 GB . It works because the routed experts stay on SSD and only get read when a token actually routes to them. Finally cleaned it up enough to publish: model: https://hugging…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-19 05:08 · r/LocalLLM
Qwen3.8 flash next, 27 tok/s — 32GB VRAM + 64GB RAM - 2026-09-18 20:15 · r/LocalLLaMA
Qwen3.8-Flash-Next (95.5 GiB) on a 64GB Mac at ~27 tok/s, checkpoint + fork