Qwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max
Hi! I'm the author of MoEspresso, which is my way of putting my own ideas about inference engines to the test. A lot of the fun has been trying different design choices, measuring what happens, and finding that several of them work well together. MoEspresso 3 runs Qwen3.8-Flash-Next on a 2021 M1 Ma…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-27 17:53 · r/LocalLLaMA
Qwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max