AINewsnow

Proper scoring rules as RL rewards: a breakdown of RLCD (the method behind TypeSafe's Jev)Standard RLVR gives +1 for a correct answer and 0 otherwise. A lucky guess and a confident correct answer get the same reward, so there's no pressure to be calibrated and the policy drifts toward overconfidence

Standard RLVR gives +1 for a correct answer and 0 otherwise. A lucky guess and a confident correct answer get the same reward, so there's no pressure to be calibrated, and the policy drifts toward overconfidence. TypeSafe's RLCD claims to fix this by rewarding calibrated probabilities. They haven't…

Read the full story at r/reinforcementlearning ↗

Timeline · 1 report

  1. 2026-09-28 11:36 · r/reinforcementlearning
    Proper scoring rules as RL rewards: a breakdown of RLCD (the method behind TypeSafe's Jev)Standard RLVR gives +1 for a correct answer and 0 otherwise. A lucky guess and a confident correct answer get the same reward, so there's no pressure to be calibrated and the policy drifts toward overconfidence

More stories

  1. OpenDecider: distilling calibrated "System One" decision models from open teachers, evaluated head-to-head against Laya and TypeSafe's Jev — r/machinelearningnews
  2. TensorSharp Jev requests can now combine documents, images, video, and audio — r/LocalLLM
  3. Jev vs. Kev: open-source Jev alternative tested side by side — r/LocalLLaMA
  4. How JEV works internally — r/learnmachinelearning
  5. New frontier AI models, TypeSafe’s Jev AI, & NASA’s IBM collab — Mixture of Experts (IBM)
  6. How long until OpenAI release their version of Jev? — r/AI_Agents
  7. Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU — MarkTechPost
  8. Jev for beginners: how to use it and what to build — How I AI

Get the daily brief of stories like this at 6:30 every morning →