llama-server's logprobs are placeholders when speculative decoding is on
This story is from 2026-10-06. It is preserved in the archive; the latest stories are on the live feed.
TL;DR : When llama-server runs with speculative decoding (a draft model with -md , MTP, or one of the n-gram types), every token that comes out of the speculative loop is sent with logprob: 0.0 and an empty top_logprobs list. The code sets the probability to 1.0 with the comment // set later and ne…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-06 03:37 · DEV Community — AI
llama-server's logprobs are placeholders when speculative decoding is on