AINewsnow

Same Claude. Different Harness. Very Different Result.

This story is from 2026-09-18. It is preserved in the archive; the latest stories are on the live feed.

Claude didn’t get smarter. We changed everything around it. Somehow, 6.5 points appeared between them. We beat Claude Code with Claude. Which is a slightly ridiculous sentence, but it is also a useful one. We ran Backboard CLI on Terminal-Bench 2.1 using Claude Opus 4.8 through Amazon Bedrock. Clau…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-18 15:20 · DEV Community — AI
    Same Claude. Different Harness. Very Different Result.

More stories

  1. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI
  2. Bolt Adds DeepSeek V4.1 Flash at 10x Cheaper Than V4 Pro — AlphaSignal
  3. Claude, Anthropic’s AI model, is helping to develop the next version of itself — Fast Company AI
  4. AI skills — r/AI_Agents
  5. OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot — The Guardian AI
  6. How To Use Ai and create those videos — r/aivideo
  7. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  8. Hitting usage limits on Codex and Claude Code.. which paid plans or setups give the most usable capacity for the money? — r/ChatGPTCoding

Get the daily brief of stories like this at 6:30 every morning →