How are you all actually evaluating agent decisions, not just agent outputs?
This story is from 2026-08-21. It is preserved in the archive; the latest stories are on the live feed.
Most agent eval I see (DeepEval, faithfulness scoring, etc) checks whether the OUTPUT is good — is it faithful, did it resist a prompt injection, etc. Pass/fail. But I've been building an agent that makes an actual decision with a cost attached (pay a supplier / verify / escalate), and pass/fail fe…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-08-21 15:43 · r/AI_Agents
How are you all actually evaluating agent decisions, not just agent outputs?