AINewsnow

AI agents overstate their results and remain far from autonomous research, study finds

Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came from methods resear…

Read the full story at The Decoder ↗

Timeline · 1 report

  1. 2026-10-11 13:44 · The Decoder
    AI agents overstate their results and remain far from autonomous research, study finds

More stories

  1. Usage resets... — r/OpenAI
  2. I Feel Like Using Cluade is Paying to Constantly Be Told No — r/ClaudeAI
  3. What Should I Do? — r/ChatGPTPro
  4. Codex (gpt-6-sol high) Comportamento Evasivo e Confissões Alucinatórias: Um Estudo de Caso sobre Continuidade de Sessão de Agentes e Proveniência — r/PromptEngineering
  5. Anthropic needs a way to change your email — r/ClaudeAI
  6. New to Claude after years of ChatGPT/Codex. What setup actually works for you? — r/learnmachinelearning
  7. Deal Days gets you ChatGPT, Claude, Gemini, and more for just $59.97 for life — Mashable AI
  8. I ran out of Claude Code tokens and had to finish my project with ChatGPT. I was genuinely surprised. — r/artificial

Get the daily brief of stories like this at 6:30 every morning →