Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2608.24814v1 Announce Type: cross Abstract: We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse thr…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-02 04:00 · arXiv stat.ML
Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining