AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
arXiv:2610.11050v1 Announce Type: new Abstract: Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spann…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.AI
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks