AINewsnow

A benchmark caught models inventing 70.7% of missing fields. Coding agents have the same failure.

This story is from 2026-09-28. It is preserved in the archive; the latest stories are on the live feed.

There's a benchmark that answers a question I keep getting asked in different clothes: when the value isn't there, what does the model do? The setup is at earnanhonestdollar.com/bench , and it's one of the cleaner eval designs I've read this month. The question underneath it: an agent buying a serv…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-28 17:45 · DEV Community — AI
    A benchmark caught models inventing 70.7% of missing fields. Coding agents have the same failure.

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  3. Meta Taps MongoDB CEO to Lead New Enterprise AI Platform — Bloomberg AI
  4. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  5. Scoop: Anthropic's Dario Amodei to have White House dinner with Trump — Axios AI+
  6. Bill Gates says unchecked AI could ‘cause a billion deaths’ in call for regulation — The Guardian AI
  7. OpenAI agents posted user images online, disclose dozens of third party incidents — Axios AI+
  8. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times Technology

Get the daily brief of stories like this at 6:30 every morning →