Jev vs LLM judges on my agent's evals: 180x cheaper, 2.6 points less accurate
I wanted to know if Jev could replace the LLM judges (Sonnet, Haiku, Gemini Flash Lite) in my agent's eval suite, so I replayed 193 real eval verdicts on Jev and three LLM judges, scored against hand labels of the verdicts, 5 runs each. Jev: 90.6% at $0.07 per 1k verdicts Sonnet 4.6: 93.2% at $13.2…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-10-10 15:25 · r/AI_Agents
Jev vs LLM judges on my agent's evals: 180x cheaper, 2.6 points less accurate