Running Qwen3.8-Flash-Next on a 128 GB Mac: The Expert-Pruning Trap, and a Memory-Mapped n-gram Table That Gets You to 240K Tokens
Hello, everyone. Have you ever wanted to run the smartest model you can on your own Mac? I have. The catch is that the smartest models are also the biggest, and even 128 GB of memory often falls just short. That is exactly when a smaller, pruned build starts to look tempting. Today's story is about…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-24 13:06 · DEV Community — Machine Learning
Running Qwen3.8-Flash-Next on a 128 GB Mac: The Expert-Pruning Trap, and a Memory-Mapped n-gram Table That Gets You to 240K Tokens