AINewsnow

Two "Codex CLI" models on the same benchmark: the harness hides the model

This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.

Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name. Same harness, two different brains: GPT-6 Astra on Codex CLI: 33.8% resolution GPT-5.6 Sol on Codex CLI: 16.2% resolution…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-14 00:15 · DEV Community — AI
    Two "Codex CLI" models on the same benchmark: the harness hides the model

More stories

  1. OpenAI discloses six new safety incidents — Axios AI+
  2. Ai agent template — r/ArtificialInteligence
  3. Dumbest solution to the alignment problem — r/singularity
  4. I spent hours going through 100+ page PDFs, so I built a tool that highlights exactly where the answer came from. It's now completely open-source. — r/ChatGPTCoding
  5. AI cybersecurity risks explode as Claude used to break into ChatGPT — Semafor Technology
  6. The cloud outage that should terrify the CIO — InfoWorld AI
  7. ChatGPT co-creator launches a new kind of AI — The Rundown AI
  8. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI

Get the daily brief of stories like this at 6:30 every morning →