Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks
arXiv:2609.21039v1 Announce Type: new Abstract: A pervasive structural pattern in modern deep learning is the linear factorization block: a submodule of the form $W = BA$ in which two parameter matrices are multiplied directly, with no intervening nonlinearity. Such blocks appear in LoRA adapters,…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-21 04:00 · arXiv cs.LG
Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks