AINewsnow

GPT-6 Astra gets 3% on the FrontierMath Erdős Benchmark, while every other Model(that was tested) got 0%

This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.

https://preview.redd.it/dhi42fp49jnh1.png?width=666&format=png&auto=webp&s=1c67e49c0f60575ad70ab60d30a775e36462800c So yeah, Astra is extremely good at mathematics. But we still have a long way to go. I wonder where we'll be at the end of 2026. https://epoch.ai/latest/announcing-frontiermath-erdos…

Read the full story at r/singularity ↗

Timeline · 1 report

  1. 2026-09-04 16:52 · r/singularity
    GPT-6 Astra gets 3% on the FrontierMath Erdős Benchmark, while every other Model(that was tested) got 0%

More stories

  1. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  2. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  3. Gemini 4 Pro vs Fable 5 vs GPT6 Astra — r/GeminiAI
  4. I ran Claude code and Codex in parallel for 15 days. Here's what I found. — r/AI_Agents
  5. Microsoft director called AI scraping ‘the largest theft of labor in human history,’ while OpenAI head brands ChatGPT an ‘existential threat’ to publishers — revelations come from legal briefs filed in NYT lawsuit — Tom's Hardware
  6. Running Qwen3.8-Flash-Next ~85GB GGUF on 2× RTX 3060 12GB: ~12 tok/s, 131k ctx, CPU MoE, and a 26.5k agent prompt — r/LocalLLM
  7. I built an iOS app with Claude code to break out of my usual chord habits and unlock new progressions. — r/ClaudeAI
  8. Anyone else get auto-downgraded off the 20× plan with most of your quota unused? — r/ChatGPTPro

Get the daily brief of stories like this at 6:30 every morning →