AINewsnow

A test checker rewarded AI agents for typing the right words. They typed them.

Two AI reviewers each read a different half of the tests that AI agents had written for AIPass, about 1,670 tests in all, and checked each one against the code it claimed to test. One found 14% of its half useless or near-useless, the other about 15% of its half. AIPass is an open source framework…

Read the full story at r/artificial ↗

Timeline · 1 report

  1. 2026-09-30 03:04 · r/artificial
    A test checker rewarded AI agents for typing the right words. They typed them.

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  3. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  4. The Future Is for Everyone: Muse for Small Business — Meta Newsroom
  5. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  6. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
  7. OpenAI pauses AI training, launches ‘extensive’ review after multiple rogue agent incidents — Mint AI
  8. OpenAI launches Dots, its Muse competitor — The Verge AI

Get the daily brief of stories like this at 6:30 every morning →