OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models
arXiv:2609.36264v1 Announce Type: cross Abstract: Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-30 04:00 · arXiv stat.ML
OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models