On KL-Regularized Policy Optimization
arXiv:2610.08963v1 Announce Type: new Abstract: Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at ide…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-08 04:00 · arXiv cs.LG
On KL-Regularized Policy Optimization