Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents
arXiv:2609.22120v1 Announce Type: new Abstract: Test-time self-evolving agents improve by reusing past experience, yet sparse-reward trajectories contain failures, loops, and detours, while summaries often omit the state conditions and action dependencies needed for execution. We study executable W…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.LG
Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents