Don't Repeat Yourself: Self-Supervised Fine-Tuning for Coverage
arXiv:2609.31688v1 Announce Type: new Abstract: In verifiable domains such as math and coding, finding one correct solution among many attempts can matter more than the pass rate of each attempt. Post-training can concentrate large language model outputs around a few modes, while increasing samplin…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv cs.CL
Don't Repeat Yourself: Self-Supervised Fine-Tuning for Coverage