AINewsnow

Frontier AI models are beginning to cross the human baseline on SimpleBench

This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.

​ EDIT: GPT-6 Astra hasn't been measured yet. They usually take a couple of days since release. u/Alt_Restorer also pointed out "the benchmark is private, so Phillip only runs it if there's an API that doesn't retain data. I'm not sure whether Astra has that yet." Context for those unfamilia…

Read the full story at r/singularity ↗

Timeline · 1 report

  1. 2026-09-06 01:58 · r/singularity
    Frontier AI models are beginning to cross the human baseline on SimpleBench

More stories

  1. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  2. Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History — New York Times Technology
  3. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  4. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  5. what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn? — r/LocalLLaMA
  6. I ran Claude code and Codex in parallel for 15 days. Here's what I found. — r/AI_Agents
  7. ChatGPT for Word is now available — OpenAI YouTube
  8. Jump Trading points GPT-6 Astra to its most ambiguous, difficult tasks — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →