Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered
This story is from 2026-09-29. It is preserved in the archive; the latest stories are on the live feed.
TL;DR — Darwin-180B-RSI, an open-weight model from Korean startup VIDRAFT, reports 100% on AIME 2026 (30/30) and 100% on HMMT February 2026 (33/33) under 16-sample majority vote with a 131K-token thinking budget. They are the first perfect scores on these two Hugging Face official leaderboards. The…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-29 12:19 · DEV Community — Machine Learning
Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered