MOE models bigger than Qwen 3.6 35B but smaller than GLM 4.5?
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
So I've got a setup that's just a medium end laptop (16 GB ddr5, intel 13th i5, no GPU, 1TB ssd) and after hours of tricks and optimisation I've managed to get Qwen 3.6 35B running comfortably on terminal (llama.cpp) at ~15 output tokens per second. The thing is I don't need that speed since I don'…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-15 13:18 · r/LocalLLM
MOE models bigger than Qwen 3.6 35B but smaller than GLM 4.5?