AINewsnow

[R] A preregistered test of TypeSafe Jev's calibration under human disagreement (ChaosNLI, 100 labels per item)

Paste the "short version" and "what I did" sections of the written post, then the limits and disclosure, then: Paper: https://zenodo.org/records/22971492 Preregistration: https://zenodo.org/records/22971413 Code: https://github.com/GautamTalksDev/jevbench

Read the full story at r/learnmachinelearning ↗

Timeline · 1 report

  1. 2026-09-29 04:51 · r/learnmachinelearning
    [R] A preregistered test of TypeSafe Jev's calibration under human disagreement (ChaosNLI, 100 labels per item)

More stories

  1. Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights) — r/LocalLLaMA
  2. OpenDecider: distilling calibrated "System One" decision models from open teachers, evaluated head-to-head against Laya and TypeSafe's Jev — r/machinelearningnews
  3. TensorSharp Jev requests can now combine documents, images, video, and audio — r/LocalLLM
  4. New frontier AI models, TypeSafe’s Jev AI, & NASA’s IBM collab — Mixture of Experts (IBM)
  5. LLM or JEV? Why not both? - introducing a hybrid Gemma4 approach — r/LocalLLM
  6. How long until OpenAI release their version of Jev? — r/AI_Agents
  7. Benchmarked Solar Decide, 57.8% vs Jev 83.3% — r/ArtificialInteligence
  8. Jev won't replace security engineers, apparently — r/PromptEngineering

Get the daily brief of stories like this at 6:30 every morning →