Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations
arXiv:2609.33186v1 Announce Type: new Abstract: Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of on…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv stat.ML
Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations