Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model
arXiv:2610.06931v1 Announce Type: new Abstract: In this paper, we study the sample complexities of value and policy learning in finite discounted Markov decision processes (MDPs) under recursive entropic risk preferences with risk parameter \(\beta\neq 0\), assuming access to a generative model of…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.LG
Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model