When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO
arXiv:2610.06861v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \emph{external guidance…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.LG
When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO