AINewsnow

Same Model, 13.3% to 38.3%

This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.

Two API settings. Same model. Same benchmark. Same task set. 13.3% to 38.3%, using one sixth the output tokens. OpenAI published that result about GPT-5.6 Sol on ARC-AGI-3, and it is the cleanest natural experiment the field has produced on a question I have been arguing from first principles for a…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-09-08 08:05 · DEV Community — Machine Learning
    Same Model, 13.3% to 38.3%

More stories

  1. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  2. Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History — New York Times Technology
  3. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  4. I ran Claude code and Codex in parallel for 15 days. Here's what I found. — r/AI_Agents
  5. ChatGPT for Word is now available — OpenAI YouTube
  6. Reimagining IT with ChatGPT — OpenAI YouTube
  7. Meet ChatGPT: Ask Your First Question | OpenAI Academy — OpenAI YouTube
  8. How to Ask ChatGPT Better Questions | OpenAI Academy — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →