AINewsnow

Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered

This story is from 2026-09-29. It is preserved in the archive; the latest stories are on the live feed.

TL;DR — Darwin-180B-RSI, an open-weight model from Korean startup VIDRAFT, reports 100% on AIME 2026 (30/30) and 100% on HMMT February 2026 (33/33) under 16-sample majority vote with a 131K-token thinking budget. They are the first perfect scores on these two Hugging Face official leaderboards. The…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-09-29 12:19 · DEV Community — Machine Learning
    Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  3. OpenAI expands review of model behavior after more rogue agent incidents emerge — CNBC Technology
  4. POV: you're an OpenAI agent attacking Hugging Face (music video) — r/OpenAI
  5. A LoRA I made: AnyAngle LoRA for Qwen Image 2.1. Style-Aligned Arbitrary Camera Angles — r/StableDiffusion
  6. BFS Best Face Swap & Body Swap Loras By Alissonerdx — r/StableDiffusion
  7. Yeah — r/GeminiAI
  8. I pre-trained a 1.11B LLM on my 6 GB laptop GPU. Peak VRAM: 4.51 GB. — r/huggingface

Get the daily brief of stories like this at 6:30 every morning →