Let the same model write the tests and the code. The tests rejected a known-correct solution 77% of the time
Classic agent setup, I built it by the book: one call writes a contract, one writes tests, one writes the code, a script runs the tests against the code, a repair step patches whatever fails. Feels like engineering. Then I did the boring check nobody does. Fed every generated test suite a known-cor…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-10-04 19:32 · r/AI_Agents
Let the same model write the tests and the code. The tests rejected a known-correct solution 77% of the time