AINewsnow

New Benchmark to see if Opus 5.5 gets nerfed

There’s been a lot of talk about Anthropic nerfing models over time. So I decided to test it. I’ll be running a series of deterministic tests on Opus 5.5 for the next month every day and will log any discrepancies. I’ll be announcing results daily on the repo: ⭐️ https://github.com/ninjahawk/livene…

Read the full story at r/ClaudeAI ↗

Timeline · 1 report

  1. 2026-09-22 23:15 · r/ClaudeAI
    New Benchmark to see if Opus 5.5 gets nerfed

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity — The Verge AI
  3. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  4. Claude Opus 5.5 now available on AI Gateway — Vercel Blog
  5. Introducing Claude Opus 5.5, 40% Cheaper and Smarter Than Ever Before — r/singularity
  6. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  7. Anthropic and OpenAI roll out cheaper models in first release since call for slowdown — CNBC Technology
  8. Claude Opus 5.5 delivers Fable 5.1 performance – and costs 40% less — ZDNET AI

Get the daily brief of stories like this at 6:30 every morning →