Speculative Evaluation of Stochastic LLMs
arXiv:2609.28560v1 Announce Type: new Abstract: Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of a…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv stat.ML
Speculative Evaluation of Stochastic LLMs