You can now run Qwen3.8-27B on a 2021 M1 Max at 39 tok/s: I ported Splash (M3+ only) and wrote custom Metal kernels
You can now run Qwen3.8-27B on a 2021 M1 Max at 39 tok/s: I ported Splash (M3+ only) and wrote custom Metal kernels TL;DR: 39 tok/s is the average of npanj's five-prompt Splash benchmark (short prompts, default reasoning), up from 19 tok/s with Splash's own kernels on the same Mac. In a real coding…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-24 02:56 · r/LocalLLM
You can now run Qwen3.8-27B on a 2021 M1 Max at 39 tok/s: I ported Splash (M3+ only) and wrote custom Metal kernels