Exploration-Preserving Policy Optimization
arXiv:2610.04011v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded response…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.AI
Exploration-Preserving Policy Optimization