Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
I believe I may currently hold the record for memory constrained inference for Qwen3.8–Flash-Next on Apple Silicon — needing only about 21GB of allocations. Introducing Cherenkov , an inference engine for Apple Silicon combining predictive expert streaming with optional mixed-precision execution. I…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-10 21:31 · r/LocalLLaMA
Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air