How are you actually verifying the model that ran a tool call?
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
We kept getting agent traces that said we ran claude-haiku-4-5, then the bill and the tool schema didn't match. Something downstream had swapped it. Evals still green because the JSON looked like a tool call. What we ended up doing on Conifer: a named catalog id is that id, or a typed error. Includ…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-08-26 19:17 · r/AI_Agents
How are you actually verifying the model that ran a tool call?