Proper scoring rules as RL rewards: a breakdown of RLCD (the method behind TypeSafe's Jev)Standard RLVR gives +1 for a correct answer and 0 otherwise. A lucky guess and a confident correct answer get the same reward, so there's no pressure to be calibrated and the policy drifts toward overconfidence
Standard RLVR gives +1 for a correct answer and 0 otherwise. A lucky guess and a confident correct answer get the same reward, so there's no pressure to be calibrated, and the policy drifts toward overconfidence. TypeSafe's RLCD claims to fix this by rewarding calibrated probabilities. They haven't…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-09-28 11:36 · r/reinforcementlearning
Proper scoring rules as RL rewards: a breakdown of RLCD (the method behind TypeSafe's Jev)Standard RLVR gives +1 for a correct answer and 0 otherwise. A lucky guess and a confident correct answer get the same reward, so there's no pressure to be calibrated and the policy drifts toward overconfidence