Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivit
arXiv:2609.38659v1 Announce Type: new Abstract: We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previ…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv stat.ML
Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivit