AINewsnow

[Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"

https://preview.redd.it/nulsv53o8vqh1.png?width=4500&format=png&auto=webp&s=74765dbd409f4c221640f9f6000a685f6fdbb242 Spent weekend benchmarking the Splash engine (by Incoai) and extending its architecture to native 8-bit on Apple Silicon (M5 Pro, 64 GB unified memory). Splash is a compiled C++ and…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-09-21 12:34 · r/LocalLLM
    [Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"
  2. 2026-09-21 12:25 · r/LocalLLaMA
    [Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"

More stories

  1. Apple Mac Studio (M5 Ultra) review: Huge AI and graphics power at a huge premium — Engadget
  2. iPhone owners can now submit claims in Apple’s $250 million Siri AI settlement — The Verge AI
  3. Google Takes on Apple, Microsoft With AI-Powered Laptops — Bloomberg AI
  4. Why is there a waitlist for Siri AI in iOS 27? — Engadget
  5. Q&A with Mark Gurman on the iPhone Duo, breaking Apple news, Apple's AI-native devices, Tim Cook staying as executive chair, John Ternus, Johny Srouji, and more (Nilay Patel/The Verge) — Techmeme
  6. Can John Ternus find Apple’s next big thing? — The Verge AI
  7. Got an Android Phone? Google Thinks You’ll Probably Want a Googlebook Laptop — Wired AI
  8. Gemini Joins the Hacker Club — Wall Street Journal Technology

Get the daily brief of stories like this at 6:30 every morning →