Recursive Self-Improvement via On-Policy Distillation for Reasoning
arXiv:2609.30652v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-28 04:00 · arXiv cs.CL
Recursive Self-Improvement via On-Policy Distillation for Reasoning