Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data
arXiv:2609.38880v1 Announce Type: new Abstract: We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $\gamma$, write $H=(1-\gamma)^{-1}$, and let $\mu_{\min}$ and $t_{\operatorname{mix}}$ deno…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv stat.ML
Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data