Can Representation Learning Decouple from Loss Minimization? Polar Updates Have an Answer
arXiv:2609.36240v1 Announce Type: cross Abstract: Does representation learning stop when the training loss stops improving? We study this question for matrix Muon, whose polar-normalised updates have a step length set by the gradient's rank rather than its norm. Near the edge of stability, full-bat…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-30 04:00 · arXiv stat.ML
Can Representation Learning Decouple from Loss Minimization? Polar Updates Have an Answer