Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret
arXiv:2610.07737v1 Announce Type: new Abstract: We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards. Existing work defines Nash regret as $\mathrm{NR}_T = \mu^\star - (\prod_{t=1}^T \mathbb{E}\mu_{…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv stat.ML
Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret