When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
arXiv:2610.00372v1 Announce Type: new Abstract: Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The sa…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-02 04:00 · arXiv cs.AI
When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents