Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining
arXiv:2609.25482v1 Announce Type: new Abstract: Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed model (e.g. the raw…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-23 04:00 · arXiv cs.LG
Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining