AINewsnow

Fitting one score across 40 sparse benchmarks with item response theory

This story is from 2026-10-09. It is preserved in the archive; the latest stories are on the live feed.

Part 1 described how our averaged composite failed: a model measured on three saturated legacy boards outranked flagships measured across eighteen, because hard boards lower a mean and easy boards raise it. Minimum counts and per-board z-scores did not fix it, so we changed the model. This part cov…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-09 12:39 · DEV Community — Machine Learning
    Fitting one score across 40 sparse benchmarks with item response theory

More stories

  1. GPT-6 and Intelligent UI for everyone — OpenAI News
  2. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  3. OpenAI Decisions API now available on AI Gateway — Vercel Blog
  4. Anthropic bans ‘abusive or cruel behavior’ toward Claude — The Verge AI
  5. Introducing Playground: Create and play custom games — Google AI Blog
  6. Fired OpenAI safety researchers dispute their dismissals in open letter — Engadget
  7. Sophos cuts threat investigation time by 96% with OpenAI Daybreak — OpenAI News
  8. Grok Imagine Video 1.5 Lite on AI Gateway — Vercel Blog

Get the daily brief of stories like this at 6:30 every morning →