AINewsnow

Jev vs LLM judges on my agent's evals: 180x cheaper, 2.6 points less accurate

I wanted to know if Jev could replace the LLM judges (Sonnet, Haiku, Gemini Flash Lite) in my agent's eval suite, so I replayed 193 real eval verdicts on Jev and three LLM judges, scored against hand labels of the verdicts, 5 runs each. Jev: 90.6% at $0.07 per 1k verdicts Sonnet 4.6: 93.2% at $13.2…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-10-10 15:25 · r/AI_Agents
    Jev vs LLM judges on my agent's evals: 180x cheaper, 2.6 points less accurate

More stories

  1. Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog
  2. Google Cloud introduces Gemini agent to change enterprise work — SiliconANGLE AI
  3. Gemini 4 Argon — r/GeminiAI
  4. Is Gemini Pro model down? — r/GeminiAI
  5. Anyone else having issues with the ChatGPT app? — r/ChatGPT
  6. H2O-Lightning-4B: Apache-2.0 4B Decision model, official #1 open model on JevBench (above Jev) — r/LocalLLaMA
  7. Whatever happened to BABA is AI from 2024? [D] — r/MachineLearning
  8. Gemini 4 is coming today (or tmr depending on your time zone.) — r/GeminiAI

Get the daily brief of stories like this at 6:30 every morning →