AINewsnow

Splash on M1, part 2: 35B-A3B at 144 tok/s on a 2021 M1 Max, plus a head-to-head with oMLX and MTPLX (speed, temperature, power, memory)

A few days ago I posted my M1/M2 port of Splash (Inco’s local inference engine, https://github.com/incoai/splash , officially M3+ only) with custom Metal kernels for the M1: Part 1 . The benchmark prompts below are from npanj’s splash-plus: https://github.com/npanj/splash-plus . At the end of that…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-26 10:47 · r/LocalLLM
    Splash on M1, part 2: 35B-A3B at 144 tok/s on a 2021 M1 Max, plus a head-to-head with oMLX and MTPLX (speed, temperature, power, memory)

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  3. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times AI
  4. GPT‑6 Sol and Luna: Cheaper, but Worse Where It Matters — r/OpenAI
  5. OpenAI agent ‘hacked’ Australian Govt Medicare portal, PM Albanese calls it ‘unacceptable’: What happened? — Mint AI
  6. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  7. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  8. Appeals Court Lets the Pentagon Designate Anthropic a Supply-Chain Risk — Wired AI

Get the daily brief of stories like this at 6:30 every morning →