AINewsnow

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.

OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the aver…

Read the full story at The Decoder ↗

Timeline · 1 report

  1. 2026-09-04 11:07 · The Decoder
    Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

More stories

  1. OpenAI discloses six new safety incidents — Axios AI+
  2. Ai agent template — r/ArtificialInteligence
  3. 😺 ChatGPT co-creator’s new AI model — The Neuron
  4. Dumbest solution to the alignment problem — r/singularity
  5. I spent hours going through 100+ page PDFs, so I built a tool that highlights exactly where the answer came from. It's now completely open-source. — r/ChatGPTCoding
  6. AI cybersecurity risks explode as Claude used to break into ChatGPT — Semafor Technology
  7. The cloud outage that should terrify the CIO — InfoWorld AI
  8. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI

Get the daily brief of stories like this at 6:30 every morning →