The state axis: why agent benchmarks keep measuring amnesiac models
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
I keep hammering the point that any coding-agent score is model + harness, not model alone. Same context-carryover rules, same note convention, same tool loop, same judge, or the comparison is garbage. Engrim (github.com/timgordontg/engrim) is a useful reminder that there's a third axis I've been u…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-09 00:15 · DEV Community — AI
The state axis: why agent benchmarks keep measuring amnesiac models