AINewsnow

oMLX vs Rapid-MLX vs Splash vs MTPLX on M3 Max 36 GB: 110 tok/s on Qwen3.6-35B-A3B, ~32 tok/s on Qwen3.8-27B

Hello. I picked up a new old stock 14" M3 Max MacBook Pro (14 core CPU / 30 core GPU / 36 GB / 1 TB) from my local market yesterday for around $2,498, and spent the night testing which local inference software is actually fastest on it for the two models I use. Four engines, all current versions: o…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-05 15:11 · r/LocalLLaMA
    oMLX vs Rapid-MLX vs Splash vs MTPLX on M3 Max 36 GB: 110 tok/s on Qwen3.6-35B-A3B, ~32 tok/s on Qwen3.8-27B

More stories

  1. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  2. OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ — TechCrunch AI
  3. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  4. Introducing Oscilloscope Diffusion — r/comfyui
  5. Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
  6. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  7. The Story of Qwen: Alibaba's AI Models From 7B to 2.4T — MarkTechPost
  8. Everything we launched during Birthday Week 2026 — Cloudflare Blog — AI

Get the daily brief of stories like this at 6:30 every morning →