AINewsnow

Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air

This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.

I believe I may currently hold the record for memory constrained inference for Qwen3.8–Flash-Next on Apple Silicon — needing only about 21GB of allocations. Introducing Cherenkov , an inference engine for Apple Silicon combining predictive expert streaming with optional mixed-precision execution. I…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-10 21:31 · r/LocalLLaMA
    Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air

More stories

  1. [Release] Nirvana Code: A single-binary Rust LLM engine built from the metal up for Apple Silicon (Metal 3, Persistent Prefix Cache, Speculative Decoding, Dual GGUF + MLX) — r/LocalLLM
  2. He’s the Face of AI Doomsday Fears — Wall Street Journal Technology
  3. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial
  4. Running ACE-Step 1.5 and YuE2-3B on one GPU behind a single local UI (CUDA + Apple Silicon): notes from building it — r/LocalLLM
  5. Best open-source model for an M2 Max 32GB and what closed model does it actually compare to? — r/LocalLLM
  6. Apple M6 Pro Geekbench 7 — r/LocalLLaMA
  7. Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro — r/LocalLLaMA
  8. Apple M5 Ultra Scores Big GPU Gains in Leaked Geekbench Benchmark — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →