Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
This story is from 2026-09-18. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.19170v1 Announce Type: new Abstract: Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize sampled dynamics. We construct an ergodic two-state counterexample in which…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-18 04:00 · arXiv cs.AI
Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes