Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents
arXiv:2609.31995v1 Announce Type: new Abstract: Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv cs.CL
Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents