AINewsnow

Qwen3.8-27B MTP quants on Apple M5 Max — which one is actually worth it?

This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.

Ran a full oQ2e → oQ8e sweep yesterday. Here's what the numbers say: | Quant | Speed | Quality Score | |-----------|-------------|-------------------------| | oQ2e-mtp | 44.8 tok/s | 16.9 ❌ (unusable) | | oQ3e-mtp | 38.7 tok/s | 85.2 ✅ | | oQ4e-mtp | 36.6 tok/s | 86.8 ✅ best balance | | oQ6e-mtp |…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-08-31 15:07 · r/LocalLLaMA
    Qwen3.8-27B MTP quants on Apple M5 Max — which one is actually worth it?

More stories

  1. [Release] Nirvana Code: A single-binary Rust LLM engine built from the metal up for Apple Silicon (Metal 3, Persistent Prefix Cache, Speculative Decoding, Dual GGUF + MLX) — r/LocalLLM
  2. He’s the Face of AI Doomsday Fears — Wall Street Journal Technology
  3. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial
  4. Muse Hit #2 Doing Phone Calls. Then It Asked for Your Bank. — DEV Community — AI
  5. Running ACE-Step 1.5 and YuE2-3B on one GPU behind a single local UI (CUDA + Apple Silicon): notes from building it — r/LocalLLM
  6. Best open-source model for an M2 Max 32GB and what closed model does it actually compare to? — r/LocalLLM
  7. Apple M6 Pro Geekbench 7 — r/LocalLLaMA
  8. Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →