Running Qwen 125B on Low-Spec Macs via SSD Streaming: Seeking Help with Perf & Prefill Optimization
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Hi everyone, I’ve been working on a project to run large LLMs on low-spec Mac machines (like the base M4 Mac Mini) by streaming model weights via SSD for each token. I’ve made significant progress and wanted to share my work and ask for community assistance. The Project: • Repo: https://github.com/…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-14 17:18 · r/LocalLLM
Running Qwen 125B on Low-Spec Macs via SSD Streaming: Seeking Help with Perf & Prefill Optimization