Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects
arXiv:2610.07755v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation me…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv stat.ML
Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects