AINewsnow

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon. Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Requirements: M3 or newer, macOS 26.4+, 36 GB Get started with a single command: brew install incoai/t…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-19 02:11 · r/LocalLLaMA
    Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

More stories

  1. Apple’s new chief executive built up to unveiling the ideal AI device, then said it was the iPhone. — The Next Web
  2. A closer look at the upcoming Siri AI-powered home hub, a key pillar of Apple's strategy for the home; sources: Apple started cutting Fitness+ staff (Mark Gurman/Bloomberg) — Techmeme
  3. Gemini Joins the Hacker Club — Wall Street Journal Technology
  4. Apple’s Home AI Hub Details; Apple Fitness+ Layoffs and iPhone Duo Apple Pencil — Bloomberg AI
  5. ComfyUI on Apple Silicon: no MLX, no fp8, 600-second kernel builds. So I built my own launcher — a personal project I'm sharing in case it helps someone. — r/comfyui
  6. [Release] Nirvana Code: A single-binary Rust LLM engine built from the metal up for Apple Silicon (Metal 3, Persistent Prefix Cache, Speculative Decoding, Dual GGUF + MLX) — r/LocalLLM
  7. He’s the Face of AI Doomsday Fears — Wall Street Journal Technology
  8. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial

Get the daily brief of stories like this at 6:30 every morning →