AINewsnow

Small LLM Judges Approved 11% and 41% of Wrong Answers. Then I Fixed My Own Pairwise Test.

This story is from 2026-10-07. It is preserved in the archive; the latest stories are on the live feed.

Grading small open judges against deterministic oracles, plus the harness that makes most judge calls unnecessary. Scope note, first paragraph as promised: every judge here is ≤3B parameters running locally. Nothing below carries over to frontier judges. Synthetic corruptions and natural errors are…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-07 14:32 · DEV Community — Machine Learning
    Small LLM Judges Approved 11% and 41% of Wrong Answers. Then I Fixed My Own Pairwise Test.

More stories

  1. Introducing Mistral Large 4 — Mistral AI News
  2. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  3. Sharing AI progress in mathematics — OpenAI News
  4. OpenAI has dumped 722 maths papers – now it must clean up the mess — New Scientist Technology
  5. Google launches SynthID Detector, a website that lets users detect AI-generated image, video, and audio media across dozens of common file formats (Ivan Mehta/TechCrunch) — Techmeme
  6. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  7. Introducing the Decisions API — OpenAI YouTube
  8. EmbeddingGemma 2: an open, lightweight multimodal embedding model — Google DeepMind Blog

Get the daily brief of stories like this at 6:30 every morning →