MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning
arXiv:2609.31644v1 Announce Type: new Abstract: Reinforcement learning (RL) is widely used in language-model post-training to maximize rewards assigned to individual model outputs, such as scores from binary verifiers or reward models trained on human feedback. However, applications such as synthet…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv cs.LG
MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning