Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models
arXiv:2610.10623v1 Announce Type: new Abstract: Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps. Despite their promise, effective post-training of LoopLMs remains challenging. Existing approac…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.LG
Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models