GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents
arXiv:2610.08959v1 Announce Type: new Abstract: On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this g…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-08 04:00 · arXiv cs.LG
GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents