RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers
arXiv:2609.22547v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) often improves pass@1 while falling behind its base model at larger sampling budgets $k$, a crossover read as evidence that RLVR only sharpens existing capability. We identify two limits to this re…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv stat.ML
RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers