Reinforcement Learning with Decomposed Subtasks
arXiv:2609.27035v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes di…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-24 04:00 · arXiv cs.AI
Reinforcement Learning with Decomposed Subtasks