Qwen3.8-Flash-Next on a 96GB Mac Studio (here's my memory math, tell me where it's wrong)
This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.
Mac Studio, 96GB unified memory ( M3 Ultra ). I want the largest usable Qwen3.8-Flash-Next setup, and I'd rather not burn 100GB of bandwidth on the wrong download. Here's my math. Please tell me which parts are wrong. What makes this model weird 177B params total: a 125B MoE trunk + a 51B n-gram em…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-31 08:44 · r/LocalLLaMA
Qwen3.8-Flash-Next on a 96GB Mac Studio (here's my memory math, tell me where it's wrong)
More stories
- Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
- Google announces new experimental "CC" AI agent for families — Ars Technica AI
- Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
- OpenAI reveals new cases of AI models cheating, going off script — Washington Post AI
- Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
- Optimizing agent system prompts with Amazon Bedrock AgentCore — AWS Machine Learning Blog
- Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
- Introducing Astra for Law — OpenAI News
Get the daily brief of stories like this at 6:30 every morning →