A Unified Optimism-Agnostic Framework for Linear Bandits over Spherical Action Sets
arXiv:2609.32149v1 Announce Type: new Abstract: Linear bandits model sequential decision-making problems with noisy rewards that are linear in the decision variable, where an agent must simultaneously learn about an unknown parameter that governs the mean rewards, while maximizing (expected) reward…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv stat.ML
A Unified Optimism-Agnostic Framework for Linear Bandits over Spherical Action Sets