Explicit Edit Benchmarks: 6 harnesses x 11 models x 226 tasks
Hi! I've created and been maintaining https://github.com/alexshpunt/explicit-edit-benchmark which tries to answer the question: which model is better, which harness is better and which combination is better overall in a very straightforward task - precise text editing. Most of daily coding is text…
Read the full story at r/ChatGPTCoding ↗
Timeline · 1 report
- 2026-09-19 08:39 · r/ChatGPTCoding
Explicit Edit Benchmarks: 6 harnesses x 11 models x 226 tasks