Speculative Decoding Made My vLLM Server Slower: The Acceptance Math
I added a draft model to my vLLM server, sent one curl request, and watched the tokens fly. Then I ran the load test with 64 concurrent users and total throughput went down . Same GPU, same target model, same prompts. The only change was one flag that every blog post says is a free speedup. Specula…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-07 04:42 · DEV Community — Machine Learning
Speculative Decoding Made My vLLM Server Slower: The Acceptance Math