AINewsnow

Running Qwen 125B on Low-Spec Macs via SSD Streaming: Seeking Help with Perf & Prefill Optimization

This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.

Hi everyone, I’ve been working on a project to run large LLMs on low-spec Mac machines (like the base M4 Mac Mini) by streaming model weights via SSD for each token. I’ve made significant progress and wanted to share my work and ask for community assistance. The Project: • Repo: https://github.com/…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-14 17:18 · r/LocalLLM
    Running Qwen 125B on Low-Spec Macs via SSD Streaming: Seeking Help with Perf & Prefill Optimization

More stories

  1. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  2. Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use — r/machinelearningnews
  3. Alibaba releases Qwen-Image-2.1, a 7B open-weight model it says outperforms most closed-source models, with native transparency and up to ten reference images (Qwen) — Techmeme
  4. Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti — r/LocalLLM
  5. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  6. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion
  7. Qwen 3.8 27B running on a single RTX 5090 researches and creates a full animation using only code. — r/artificial
  8. US government website used Chinese model the FBI called "malicious" — Ars Technica AI

Get the daily brief of stories like this at 6:30 every morning →