Reminder: try probabilistic MTP if you missed it. Decode +14% on prose
https://github.com/ggml-org/llama.cpp/pull/27694 Now merged. Update your llama if you haven't done so yet. Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling. Main gain seems to be on prose generation. Tests above ran with thinking off, ngram-mod off.
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-11 09:20 · r/LocalLLaMA
Reminder: try probabilistic MTP if you missed it. Decode +14% on prose