Evaluate the evaluator, judge the judge
I have been reading about llm as a judge and trying to make a tool that can help to evaluate and figure out which model can be the best judge.. My method : I keep the dataset format: * task * rubric * ideal response * negative response The negative response doesn’t necessarily have to be incorrect.…
Read the full story at r/PromptEngineering ↗
Timeline · 1 report
- 2026-09-19 17:01 · r/PromptEngineering
Evaluate the evaluator, judge the judge