Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-05 04:00 · arXiv cs.AI
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement