Which benchmarks are still far from saturation?
This is one of those things that move insanely fast. I'm mostly interested in benchmarks where frontier models score very low (30% or less preferably). The only ones that are actively maintained that I can think of are: * RLI (top score 20%) * ProgramBench (top score 4.5%) There's also a few others…
Read the full story at r/artificial ↗
Timeline · 1 report
- 2026-09-28 01:19 · r/artificial
Which benchmarks are still far from saturation?