AINewsnow

你的代理"知道"工具没用了吗?它知道,但不会停下

This story is from 2026-10-06. It is preserved in the archive; the latest stories are on the live feed.

2026 年 10 月 5 日,arXiv cs.AI 上的一篇论文: 7 个用工具的代理,面对一个持续返回无用的检索源,97%–100% 的情况下会正确判断"这个结果没用"。然后呢?它们接着查。 这是我最近看到的最能解释当下代理系统的东西:模型判断和系统行为之间的断裂。这一周至少 7 篇新论文在从不同侧面打同一个问题,串起来看,会发现 2026 年的前沿已经不在"能不能做对",而在"能不能知道自己做不对"。这篇文章我把它们串成一条线,附可复现的机制和数字。 先搭个场景 你让一个编码代理修构建脚本。它回"已解决"。你盯着答案看了三秒,发现前提里有个物理上不可能的数值——这道题根本无解。它没有…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-06 12:39 · DEV Community — AI
    你的代理"知道"工具没用了吗?它知道,但不会停下

More stories

  1. Mistral releases Mistral Large 4, dubbed "le Chonk", a 1T-parameter open-weight model for general agentic capabilities, trained on 4,000 Grace Blackwell GPUs (Sabrina Ortiz/The Deep View) — Techmeme
  2. Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
  3. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  4. OpenAI safety leader quits, warning AI company’s culture is ‘broken’ — The Guardian AI
  5. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  6. can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
  7. Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog
  8. The Story of Qwen: Alibaba's AI Models From 7B to 2.4T — MarkTechPost

Get the daily brief of stories like this at 6:30 every morning →