Cost-Aware Best-LLM Identification using Dueling Feedback
arXiv:2609.30360v1 Announce Type: cross Abstract: Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise c…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-28 04:00 · arXiv stat.ML
Cost-Aware Best-LLM Identification using Dueling Feedback