Compact Rollback MTP: a MTP version for QWEN models for those with little vRAM
This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.
I've made a modification of llama.cpp MTP for people that want to run models like QWEN 27B on 16GB and similar setup, the focus is reducing the memory cost of MTP allowing more speed for less ctx cost. MTP Mode Maximum Draft (n) Available Context TG (t/s) Standard 2 72,192 39.53 MTP Compact Rollbac…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-31 16:58 · r/LocalLLaMA
Compact Rollback MTP: a MTP version for QWEN models for those with little vRAM