AINewsnow

An agent used up to 119 tool calls, and the benchmark still only judged its final answer

This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.

FinFIRST’s Ling-3.0-flash-Fin run averaged 23.08 reasoning rounds and 30.87 tool calls per task. One task reached 119 calls. The run produced valid outputs for 122 of 123 tasks, with one missing failure. Yet the benchmark methodology says the GLM-5.1 automated judge evaluated the final answer—not t…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-09-05 14:07 · r/AI_Agents
    An agent used up to 119 tool calls, and the benchmark still only judged its final answer

More stories

  1. Zhipu Opens GLM-5.3-FlashX Near 200 Tokens/s on ~100k Domestic Accelerators — Pandaily
  2. Qwen3.8 Max (0902) scores 45 on the Artificial Analysis Intelligence Index, up 5 points in a month and back on top of China's leaderboard, nosing out GLM-5.3 (44.9) and Kimi K3 (43.8) — r/LocalLLaMA
  3. China’s Z.ai raises revenue target 25% after US$5 billion cash injection — South China Morning Post Tech
  4. Inside ZCode (Made by GLM team): Silently Uploading Your Entire Git History to the Cloud — r/LocalLLaMA
  5. Ministral 3 3B on a Galaxy S21 relayed a conversation between Gemini and Z.ai across Chrome and Firefox — r/LocalLLaMA
  6. Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
  7. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  8. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →