AINewsnow

The AI you test in the afternoon may not be the AI you test at night, even with the same name

I asked Claude Opus 5 the same 40 questions twice on the same day, a few hours apart. The second time it looked things up about 60% more often, wrote about 50% more, and a score for how well it supports its claims with sources went from 59 to 90. Same model name in every answer. Same questions. Two…

Read the full story at r/artificial ↗

Timeline · 1 report

  1. 2026-09-23 19:10 · r/artificial
    The AI you test in the afternoon may not be the AI you test at night, even with the same name

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  3. Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity — The Verge AI
  4. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  5. Anthropic and OpenAI roll out cheaper models in first release since call for slowdown — CNBC Technology
  6. Claude Opus 5.5 now available on AI Gateway — Vercel Blog
  7. AI for coding? — r/artificial
  8. Summary of METR's predeployment evaluation of Claude Opus 5.5 — METR

Get the daily brief of stories like this at 6:30 every morning →