you can now use MTP in GLM-Air
This story is from 2026-08-23. It is preserved in the archive; the latest stories are on the live feed.
If anyone still remembers GLM-4.5-Air from last year, you can now get a nice speedup by enabling MTP in llama.cpp. It is a 106B MoE with only 12B active parameters, which makes it interesting for machines with lots of memory but limited compute, such as Strix Halo or DGX Spark. I use it on 3090s. I…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-23 20:08 · r/LocalLLaMA
you can now use MTP in GLM-Air