Anatomy of a Decode-Step Stall: Where p99 Inter-Token Latency Really Comes From
This story is from 2026-10-08. It is preserved in the archive; the latest stories are on the live feed.
The One-Line Summary: In my instrumented engine, 98% of a typical inter-token gap was the decode kernel itself, but in the tokens at or above p99 the kernel was only 22–27% of the gap — the rest was a prefill chunk, a preemption, a graph miss, a garbage-collection pause or a tokenizer call that hap…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-08 10:23 · DEV Community — Machine Learning
Anatomy of a Decode-Step Stall: Where p99 Inter-Token Latency Really Comes From