Isn't "rerun the tests until green" just grading the agent on its training set?
One agent writes tests from the spec, another writes the code in parallel, and the test agent never sees the code. Their reason was basically that otherwise the AI cheats. I take it to mean tests written after reading the code just confirm what it does, bugs included. I come from ML, so to me this…
Read the full story at r/ChatGPTCoding ↗
Timeline · 1 report
- 2026-10-02 21:26 · r/ChatGPTCoding
Isn't "rerun the tests until green" just grading the agent on its training set?