Learning from the Gap Between Pass@K and Pass@1
arXiv:2609.35793v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR). An exact verifier can also support test-time scaling by selecting a passing response from multiple samples, while other deployments use…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-30 04:00 · arXiv stat.ML
Learning from the Gap Between Pass@K and Pass@1