Pooling and Drift in Delayed Bandits
This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.01761v1 Announce Type: new Abstract: A system often has to act long before it learns whether the act worked: a recommender sees a click in seconds and a purchase in days. With $K$ actions and a delay of $d$ rounds, the best rate known for this setting is $\widetilde{O}(\sqrt{(K+d)T})$ ov…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-03 04:00 · arXiv stat.ML
Pooling and Drift in Delayed Bandits