AINewsnow

The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?

This story is from 2026-08-24. It is preserved in the archive; the latest stories are on the live feed.

On July 24, ARC Prize verified Claude Opus 5 at 30.16% on the ARC-AGI-3 public set. On August 21, NVIDIA reported the same model at 100.00 on the same set. The weights did not change. The code around them did. In between, MIT did the same thing (August 5), a group led by Impossible Research got to…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-08-24 13:38 · DEV Community — Machine Learning
    The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?

More stories

  1. I built a small proxy that lets Claude Desktop / Claude Code run on local models and NVIDIA's free API, sharing it in case it's useful — r/LocalLLM
  2. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  3. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI
  4. Claude, Anthropic’s AI model, is helping to develop the next version of itself — Fast Company AI
  5. AI skills — r/AI_Agents
  6. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  7. Anthropic selects Accenture as first embedded evaluator to help implement Amodei's slowdown proposal — CNBC Technology
  8. Bolt Adds DeepSeek V4.1 Flash at 10x Cheaper Than V4 Pro — AlphaSignal

Get the daily brief of stories like this at 6:30 every morning →