Benchmarking became easy
BENCHMAXXING: How do you know a model isn’t just bench-maxxed? I tried DeepSeek-V4.1 after seeing its leaderboard numbers and… yeah, good model, but nowhere near what I expected on my actual codebase. So I built Any-Bench: (URL: Link in Commentt) That’s kind of the issue with SWE-Bench/DeepSWE/Term…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-26 05:32 · r/AI_Agents
Benchmarking became easy