Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks
arXiv:2609.33180v1 Announce Type: new Abstract: As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and wh…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv stat.ML
Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks