AINewsnow

I ran Qwen 3.8 27B on my new MacBook Pro M5 Max and on my RTX 5090 workstation. The Mac held up way better than I expected.

This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.

I have been benchmarking Qwen 3.8 27B (Q4_K_M, llama.cpp) on my RTX 5090 workstation for a while. Last week I ran the exact same sweep on my new MacBook Pro M5 Max, 36GB, just to see how bad the gap would be. The Mac is slower in raw tokens per second. No surprise there, it is a laptop sharing memo…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-05 09:51 · r/LocalLLM
    I ran Qwen 3.8 27B on my new MacBook Pro M5 Max and on my RTX 5090 workstation. The Mac held up way better than I expected.

More stories

  1. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  4. Scaling LLM on Edge Devices: A Step-by-Step Guide — DEV Community — AI
  5. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA
  6. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  7. Post-training image models for fandom — Character.AI Blog
  8. Testing Qwen 3.8 27B running locally on a single 5090 — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →