MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?
Looking for some assistance /ideation. I am running qwen flash next q4 in my Mac mini m5 64gb. QFN doesn’t fit so this is done by having as many experts hot in cache as possible and streaming in the rest from ssd. I’m getting 17.5 tks decode and 390 tks pp. Have done a bunch of optimisations includ…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-05 10:50 · r/LocalLLaMA
MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?