vLLM Preemption: Why 1 in 50 Requests Restarts From Scratch
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Our chat endpoint had a mean latency of 1.9s and a p99 of 11.4s. Same model. Same GPU. Same prompt template. nvidia-smi showed 96% utilization and no memory pressure worth mentioning. Nothing crashed. Nothing retried. But roughly one request in fifty would stream a few tokens, freeze for six second…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-14 16:55 · DEV Community — Machine Learning
vLLM Preemption: Why 1 in 50 Requests Restarts From Scratch