vLLM vs SGLang: PagedAttention vs RadixAttention Benchmarks
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
Direct Answer: In our empirical benchmarks on multi-turn agent workflows, SGLang's RadixAttention outperformed vLLM's PagedAttention Automatic Prefix Caching (APC), slashing warm Time-To-First-Token (TTFT) from 184ms to 68ms (an 82.9% latency reduction over cold prefill) and boosting sustained thro…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-10 04:02 · DEV Community — Machine Learning
vLLM vs SGLang: PagedAttention vs RadixAttention Benchmarks