AINewsnow

MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?

Looking for some assistance /ideation. I am running qwen flash next q4 in my Mac mini m5 64gb. QFN doesn’t fit so this is done by having as many experts hot in cache as possible and streaming in the rest from ssd. I’m getting 17.5 tks decode and 390 tks pp. Have done a bunch of optimisations includ…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-05 10:50 · r/LocalLLaMA
    MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?

More stories

  1. ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5. — r/huggingface
  2. Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance — r/LocalLLM
  3. Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open — r/LocalLLaMA
  4. Worth moving on from Qwen3.6 35B A3B UD on a gaming PC? — r/LocalLLM
  5. Qwen Image 2.1 Uncensored MCP — r/StableDiffusion
  6. A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. — r/LocalLLaMA
  7. How can I achieve this style in Qwen 2.1? — r/comfyui
  8. FinVector-Market-4B: A Controlled Study of LoRA Adaptation for Structured Financial Tasks — arXiv cs.LG

Get the daily brief of stories like this at 6:30 every morning →