AINewsnow

llama-server's logprobs are placeholders when speculative decoding is on

This story is from 2026-10-06. It is preserved in the archive; the latest stories are on the live feed.

TL;DR : When llama-server runs with speculative decoding (a draft model with -md , MTP, or one of the n-gram types), every token that comes out of the speculative loop is sent with logprob: 0.0 and an empty top_logprobs list. The code sets the probability to 1.0 with the comment // set later and ne…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-06 03:37 · DEV Community — AI
    llama-server's logprobs are placeholders when speculative decoding is on

More stories

  1. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  2. LLM Inference Dashboard — r/LocalLLaMA
  3. Photon Announces $4.5M Seed Round to Help Developers Build AI Agents for iMessage and WhatsApp — AI Insider
  4. llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more — r/LocalLLaMA
  5. Llama.cpp + WebGPU = agants.html — r/AI_Agents
  6. SkyIsNotGreen/Scion-35B-A3B · Hugging Face - Ternary MoE — r/LocalLLaMA
  7. WHIRL v0.1.3 — native Windows LLM engine for the Radeon AI PRO R9700: up to 2.8× llama.cpp on the same GGUF, same answers bit-for-bit — r/LocalLLM
  8. Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →