AINewsnow

How are you catching agent tests that pass on the old code?text

This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.

Cursor and Claude keep marking tests green for me when the test never failed on the previous code. I have been reverting the PR diff by hand and rerunning only the new tests. If they still pass, they were not attached to the change. I wrapped that in a GitHub Action for pytest. Curious what other p…

Read the full story at r/ChatGPTCoding ↗

Timeline · 1 report

  1. 2026-09-16 16:21 · r/ChatGPTCoding
    How are you catching agent tests that pass on the old code?text

More stories

  1. A zero-click RCE flaw in AI coding agents could have exposed enterprise systems — InfoWorld AI
  2. I ran Claude code and Codex in parallel for 15 days. Here's what I found. — r/AI_Agents
  3. How are you actually catching unsafe stuff before an agent runs it, not after? — r/AI_Agents
  4. Best AI coding agent that understands tasks well AND doesn't drain usage limits fast? (Codex vs Cursor vs Claude) — r/AI_Agents
  5. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI
  6. Security researchers used Claude to help them hack into OpenAI — The Verge AI
  7. Claude, Anthropic’s AI model, is helping to develop the next version of itself — Fast Company AI
  8. Is this Minimax H3? — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →