CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters
arXiv:2610.00321v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding veri…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-02 04:00 · arXiv cs.CL
CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters