CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
arXiv:2609.30471v1 Announce Type: new Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct proce…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-28 04:00 · arXiv cs.CL
CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production