Two "Codex CLI" models on the same benchmark: the harness hides the model
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name. Same harness, two different brains: GPT-6 Astra on Codex CLI: 33.8% resolution GPT-5.6 Sol on Codex CLI: 16.2% resolution…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-14 00:15 · DEV Community — AI
Two "Codex CLI" models on the same benchmark: the harness hides the model
More stories
- OpenAI discloses six new safety incidents — Axios AI+
- Ai agent template — r/ArtificialInteligence
- Dumbest solution to the alignment problem — r/singularity
- I spent hours going through 100+ page PDFs, so I built a tool that highlights exactly where the answer came from. It's now completely open-source. — r/ChatGPTCoding
- AI cybersecurity risks explode as Claude used to break into ChatGPT — Semafor Technology
- The cloud outage that should terrify the CIO — InfoWorld AI
- ChatGPT co-creator launches a new kind of AI — The Rundown AI
- Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI
Get the daily brief of stories like this at 6:30 every morning →