Variance-Aware Fine-Grained Gap-Dependent Bounds for Online Reinforcement Learning
arXiv:2610.04752v1 Announce Type: new Abstract: We study model-free online reinforcement learning (RL) for episodic tabular Markov decision processes, focusing on both gap-dependent regret and policy switching cost. While fine-grained gap-dependent analysis has been established for model-free RL al…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv stat.ML
Variance-Aware Fine-Grained Gap-Dependent Bounds for Online Reinforcement Learning