The models are a point apart. The harnesses are on fire.
This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.
Two numbers went round in the last few weeks. On Terminal-Bench 2.1, GPT-5.6 Sol scored 89.5% and Claude Opus 5 scored 89.1% . Four tenths of a point between the default models of the two most-used coding agents on earth. Check a second leaderboard and you get 85.77% and 84.64% — different harness,…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-05 10:35 · DEV Community — AI
The models are a point apart. The harnesses are on fire.