FrontierHarness tested 9 agent harnesses with same model, cost per pass varies 17x
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
A year ago the question was which model. Now it's which harness. Pi, Exo, Claude Code, Codex, DeepSeek Harness and 4 others. Same model, same tasks, same runtime. 360 runs, 2 billion tokens. Pass rates: 50% to 67%. Cost per pass: $1.05 to $18.34. FrontierHarness tested 9 harnesses on 30 tasks using…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-09 19:11 · r/AI_Agents
FrontierHarness tested 9 agent harnesses with same model, cost per pass varies 17x