How Inefficient Is Natural Gradient Descent? From Exact Optimality to \Theta ( \sqrt{ \log d } ) Divergence
arXiv:2610.07228v1 Announce Type: new Abstract: Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback--Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher--Rao path. We quantify this o…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv stat.ML
How Inefficient Is Natural Gradient Descent? From Exact Optimality to \Theta ( \sqrt{ \log d } ) Divergence