AINewsnow

GPT-6 Astra scored 99.9% on the AGI benchmark. The safety findings are the real story.

This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.

Less than 24 hours ago, OpenAI released GPT-6 Astra. I have been reading everything since. This is what I found and how I feel about it. TL;DR: GPT-6 Astra scored 99.9% on ARC-AGI-3 using OpenAI's proprietary evaluation harness (62.7% on the standard independent harness). It beats Claude Fable 5.1…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-09-04 04:00 · DEV Community — Machine Learning
    GPT-6 Astra scored 99.9% on the AGI benchmark. The safety findings are the real story.

More stories

  1. OpenAI discloses six new safety incidents — Axios AI+
  2. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  3. Tested Cursor, Claude Code, Codex and Antigravity on the exact same app build — r/AI_Agents
  4. Dumbest solution to the alignment problem — r/singularity
  5. AI cybersecurity risks explode as Claude used to break into ChatGPT — Semafor Technology
  6. The cloud outage that should terrify the CIO — InfoWorld AI
  7. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI
  8. I built a free browser tool for assembling reusable AI prompts. Would you use this instead of saved prompts? — r/PromptEngineering

Get the daily brief of stories like this at 6:30 every morning →