AINewsnow

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

arXiv:2609.00038v1 Announce Type: new Abstract: Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well. The metric is structurally blind to an agent that reaches the right answer the wrong way. We measure th…

Read the full story at arXiv cs.CL ↗

Timeline · 1 report

  1. 2026-09-02 04:00 · arXiv cs.CL
    trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

More stories

  1. Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
  2. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  3. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  4. Optimizing agent system prompts with Amazon Bedrock AgentCore — AWS Machine Learning Blog
  5. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  6. Introducing Astra for Law — OpenAI News
  7. OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system — The Guardian AI
  8. What It Takes to Bring Up a Multi-Rack NVIDIA Vera Rubin NVL72 Cluster — CoreWeave Blog

Get the daily brief of stories like this at 6:30 every morning →