Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance
arXiv:2610.02355v1 Announce Type: new Abstract: Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not well understood. Analyses of stochastic optimization often assume uniformly bounded s…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-05 04:00 · arXiv cs.LG
Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance