Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depe…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.AI
Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits