AINewsnow

Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered

TL;DR — Darwin-180B-RSI, an open-weight model from Korean startup VIDRAFT, reports 100% on AIME 2026 (30/30) and 100% on HMMT February 2026 (33/33) under 16-sample majority vote with a 131K-token thinking budget. They are the first perfect scores on these two Hugging Face official leaderboards. The…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 2 reports

  1. 2026-09-28 08:03 · DEV Community — Machine Learning
    Korean startup's open 180B model tops five Hugging Face official leaderboards, including perfect AIME 2026 and HMMT 2026 scores
  2. 2026-09-28 08:01 · DEV Community — Machine Learning
    Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered

More stories

  1. PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill — r/LocalLLM
  2. Scoop: Top AI companies probing tens of thousands of security incidents — Axios AI+
  3. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  4. POV: you're an OpenAI agent attacking Hugging Face (music video) — r/OpenAI
  5. Nvidia Debuts System Designed to Stop AI Agents From Going Awry — Bloomberg AI
  6. BFS Best Face Swap & Body Swap Loras By Alissonerdx — r/StableDiffusion
  7. Yeah — r/GeminiAI
  8. I pre-trained a 1.11B LLM on my 6 GB laptop GPU. Peak VRAM: 4.51 GB. — r/huggingface

Get the daily brief of stories like this at 6:30 every morning →