AINewsnow

Reminder: try probabilistic MTP if you missed it. Decode +14% on prose

https://github.com/ggml-org/llama.cpp/pull/27694 Now merged. Update your llama if you haven't done so yet. Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling. Main gain seems to be on prose generation. Tests above ran with thinking off, ngram-mod off.

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-11 09:20 · r/LocalLLaMA
    Reminder: try probabilistic MTP if you missed it. Decode +14% on prose

More stories

  1. Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM — r/LocalLLaMA
  2. VS Code Copilot + llama-server --parallel 2: concurrent requests cause full re-prefill — r/LocalLLM
  3. Pancho The Llama (Pt I) - Mean People Suck. — r/comfyui
  4. local semantic file search for Linux (Rust, llama.cpp, EmbeddingGemma 2) — r/LocalLLaMA
  5. Nemotron 3 Super (120B) at 43 tok/s on one RTX 4090, 2.5× faster than llama.cpp — r/LocalLLM
  6. [Model] Support MiniCPM-V 4.7 by tc-mb · Pull Request #29416 · ggml-org/llama.cpp — r/LocalLLaMA
  7. Comfyui and GPU on different machines? — r/comfyui
  8. Qwen3.8-Flash-Next (125B) at ~100 tok/s on an M5 Ultra Mac Studio with llama.cpp — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →