REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement
arXiv:2609.21108v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) has achieved strong performance across a wide range of continuous-control problems. These continuous-control policies, however, are often defined as direct mappings from an observed state to an action or action distri…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-21 04:00 · arXiv cs.LG
REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement