AINewsnow

You can now run Qwen3.8-27B on a 2021 M1 Max at 39 tok/s: I ported Splash (M3+ only) and wrote custom Metal kernels

You can now run Qwen3.8-27B on a 2021 M1 Max at 39 tok/s: I ported Splash (M3+ only) and wrote custom Metal kernels TL;DR: 39 tok/s is the average of npanj's five-prompt Splash benchmark (short prompts, default reasoning), up from 19 tok/s with Splash's own kernels on the same Mac. In a real coding…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-24 02:56 · r/LocalLLM
    You can now run Qwen3.8-27B on a 2021 M1 Max at 39 tok/s: I ported Splash (M3+ only) and wrote custom Metal kernels

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Bringing Private Processing to Meta AI Glasses — Engineering at Meta
  3. Gemini 3.8 text-to-speech models now available on AI Gateway — Vercel Blog
  4. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  5. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  6. OpenAI Agent Hacked Australian Government Website — Wall Street Journal Technology
  7. AI Exchange — Financial Times AI
  8. How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →