AINewsnow

How do you count a partial edit when you're measuring how often humans undo the agent's work?

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

We gave up on task success rate a while ago. An agent can complete a task cleanly and still make the wrong call, and that counts as a success, so the number kept going up while people quietly stopped trusting it. What we use now is blunter. We diff the record 48 hours after the agent touched it, an…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-09-02 01:36 · r/AI_Agents
    How do you count a partial edit when you're measuring how often humans undo the agent's work?

More stories

  1. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  2. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  3. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  4. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  5. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  6. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  7. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  8. AI hallucination of Chinese nuclear components almost led to US military attack — Ars Technica AI

Get the daily brief of stories like this at 6:30 every morning →