Qwen3.8-Flash-Next-oQ4e-mtp: 45 tok/s on M4 Max, 25 tok/s on M2 Ultra for local inference — llm-bench.io
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
Qwen 3.8 Flash Next gives similar speed than Qwen3.8 27B on Apple Silicon.
Read the full story at r/LocalLLaMA ↗
Timeline · 4 reports
- 2026-09-09 01:34 · r/LocalLLaMA
Qwen3.8-Flash-Next on MLX-serve, 1m context is released! - 2026-09-07 10:43 · r/LocalLLM
Qwen3.8 Flash Next - Strix Halo - 2026-09-06 16:57 · r/LocalLLM
Qwen3.8 Flash Next is the best model I tested for 128Gb - 2026-09-06 08:42 · r/LocalLLaMA
Qwen3.8-Flash-Next-oQ4e-mtp: 45 tok/s on M4 Max, 25 tok/s on M2 Ultra for local inference — llm-bench.io