OSCAR: Order-aware Scoring and Calibration for AI Rankings
arXiv:2609.24128v1 Announce Type: new Abstract: Judge-specific sensitivity is useful for aggregating pairwise LLM evaluations, but its interpretation depends on which systematic presentation effects the ranking model includes. We introduce OSCAR, an order-aware framework for scoring and calibrating…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv stat.ML
OSCAR: Order-aware Scoring and Calibration for AI Rankings