On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning
arXiv:2609.26918v1 Announce Type: new Abstract: Hindsight relabeling which retroactively replacing a transition's goal with the outcome the agent actually achieved is an effective tool for improving sample-efficiency in Reinforcement Learning (RL). A natural extension to preference-conditioned mult…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-24 04:00 · arXiv cs.LG
On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning