Qwen3.8-Flash-Next on a 48GB M5 Pro MBP: ~20 tok/s in chat, but slow prefill makes agent use unpractical
This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.
I’ve been testing Unsloth’s UD-IQ1_M GGUF on my M5 Pro MacBook Pro with 48GB unified memory, using llama.cpp b10930 (commit 56381e407). Mmap initially caused Metal OOM, even with all experts on CPU. Switching to --load-mode none --lazy-mode on resolved those errors in my tests while keeping the PLE…
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-09-14 04:30 · r/LocalLLaMA
Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra - 2026-09-12 21:08 · r/LocalLLaMA
Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo - 2026-09-12 15:27 · r/LocalLLM
Qwen3.8-Flash-Next on a 48GB M5 Pro MBP: ~20 tok/s in chat, but slow prefill makes agent use unpractical