Turns out that many current science-based LLM benchmarks have flaws in their answers. When corrected, the LLM benchmark scores rose significantly.
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
Paper on Arxiv
Read the full story at r/singularity ↗