Qwen3.8-Flash-Next Open Weights Model Runs on 128GB Macs
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
Qwen released Qwen3.8-Flash-Next, a 125B-token multimodal MoE previewing Qwen4 architecture. Community benchmarks show it running at full 262K context on 128GB Macs via oMLX and llama.cpp, with early support limitations.
Read the full story at Simon Willison's Weblog ↗
Timeline · 4 reports
- 2026-08-29 13:32 · r/LocalLLM
Honey, i shrunk Qwen3. 8-Flash-Next - 2026-08-28 02:25 · r/LocalLLM
Running Qwen3.8-Flash-Next at Full 262K Context on a 128GB MacBook - 2026-08-27 12:34 · r/LocalLLaMA
Qwen3.8-Flash-Next: Time to Update Those Benchmarks - 2026-08-26 23:52 · Simon Willison's Weblog
Qwen3.8-Flash-Next