AINewsnow

21 LLM judges, 541,000 verdicts, 9 providers. Raw agreement overstated real accuracy by 33-41 points on every model.

A team ran 541,000 judgments across 21 LLM judges from nine providers. The finding that held up everywhere: exact-match agreement overstates chance-corrected accuracy by 33 to 41% points. On every single model. The paper is "Reliability without Validity" (arXiv:2606.19544, June 2026). The authors c…

Read the full story at r/PromptEngineering ↗

Timeline · 1 report

  1. 2026-09-23 14:13 · r/PromptEngineering
    21 LLM judges, 541,000 verdicts, 9 providers. Raw agreement overstated real accuracy by 33-41 points on every model.

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Gemini 3.8 text-to-speech says hello — Google Gemini Blog
  3. Google releases Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its "most expressive audio generation models yet", with support for more than 100 languages (Google) — Techmeme
  4. **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** — Hugging Face Blog
  5. Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity — The Verge AI
  6. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  7. No Shirt, No Shoes, No Service: Amazon Blocks Meta’s Muse AI From Shopping — CNET AI
  8. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech

Get the daily brief of stories like this at 6:30 every morning →