Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.02897v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-04 04:00 · arXiv cs.CL
Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding