AINewsnow

GPT-6 Astra makes a massive leap on ZeroBench (an extremely difficult vision benchmark), surpassing the human baseline across all three metrics

pass@5: Scores if at least one of the 5 attempts is correct pass^5: Scores only if all 5 attempts are correct (a reliability metric) "pass@1": Not a true single-attempt pass@1, it's the average score across 5 attempts. https://zerobench.github.io/

Read the full story at r/singularity ↗

Timeline · 1 report

  1. 2026-09-23 13:00 · r/singularity
    GPT-6 Astra makes a massive leap on ZeroBench (an extremely difficult vision benchmark), surpassing the human baseline across all three metrics

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity — The Verge AI
  3. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  4. British Columbia Sues OpenAI, Alleging ChatGPT Aided Mass School Shooting — Wall Street Journal Technology
  5. Anthropic and OpenAI roll out cheaper models in first release since call for slowdown — CNBC Technology
  6. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  7. Claude Opus 5.5 delivers Fable 5.1 performance – and costs 40% less — ZDNET AI
  8. What's your proudest side-project made with Claude? — r/ClaudeAI

Get the daily brief of stories like this at 6:30 every morning →