High-Probability Guarantees for SGD under $\beta$-Heavy-Tailed Gradient Noise
arXiv:2609.32195v1 Announce Type: new Abstract: Stochastic gradient descent (SGD) is widely used to train machine learning models, but subsampling the training data introduces noise into its updates. The strength and applicability of high-probability guarantees therefore depend critically on how th…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv stat.ML
High-Probability Guarantees for SGD under $\beta$-Heavy-Tailed Gradient Noise