Fitting one score across 40 sparse benchmarks with item response theory
This story is from 2026-10-09. It is preserved in the archive; the latest stories are on the live feed.
Part 1 described how our averaged composite failed: a model measured on three saturated legacy boards outranked flagships measured across eighteen, because hard boards lower a mean and easy boards raise it. Minimum counts and per-board z-scores did not fix it, so we changed the model. This part cov…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-09 12:39 · DEV Community — Machine Learning
Fitting one score across 40 sparse benchmarks with item response theory