One behaviour explains most of a 4-point score difference. Here's the trace.
This story is from 2026-09-29. It is preserved in the archive; the latest stories are on the live feed.
When one platform beats another across 8 tests, the useful question isn't which won. It's whether the wins share a cause. In this case most of them do. One behaviour, reproducible on two separate models, accounts for the four largest score differences in the set. The tests where that behaviour does…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-29 19:15 · DEV Community — AI
One behaviour explains most of a 4-point score difference. Here's the trace.