An agent used up to 119 tool calls, and the benchmark still only judged its final answer
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
FinFIRST’s Ling-3.0-flash-Fin run averaged 23.08 reasoning rounds and 30.87 tool calls per task. One task reached 119 calls. The run produced valid outputs for 122 of 123 tasks, with one missing failure. Yet the benchmark methodology says the GLM-5.1 automated judge evaluated the final answer—not t…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-05 14:07 · r/AI_Agents
An agent used up to 119 tool calls, and the benchmark still only judged its final answer