What Pretraining and Midtraining Make Learnable from Rewards?
arXiv:2609.38446v1 Announce Type: new Abstract: A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computa…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.LG
What Pretraining and Midtraining Make Learnable from Rewards?