Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and still 35 tok/s with 400K tokens in the context
After my Splash M1 port the most common request was Qwen3.8-Flash-Next. Adding its architecture to Splash from scratch (hybrid recurrent layers, indexed sparse attention, n-gram tables, MTP) would take forever, so I took antirez's ds4 ( DwarfStar ), which already runs it, and spent the time making…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-03 18:45 · r/LocalLLM
Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and still 35 tok/s with 400K tokens in the context
More stories
- NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
- Guided Vision in Gemini Live: built for accessibility — Google Gemini Blog
- Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
- Google unveils Gemini 4 Argon: Its most powerful AI model yet, focused on coding and cyber defence — Mint AI
- OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
- A model guide for the GPT-6 family — OpenAI News
- The latest AI news we announced in September 2026 — Google Gemini Blog
- Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
Get the daily brief of stories like this at 6:30 every morning →