Running 176b Qwen3.8-Flash-Next-MLX-oQ4 on a 16GB M4
This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.
It may be useful to someone so I created a github and decided to post here. I'm getting around 1.9 tok/s. Plan on continue research to improve the speeds and maybe do a proper GUI. https://github.com/1architect/macqwen-releases
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-28 03:11 · r/LocalLLM
Running 176b Qwen3.8-Flash-Next-MLX-oQ4 on a 16GB M4