stat.MLOct 5, 2026

How Inefficient Is Natural Gradient Descent? From Exact Optimality to Θ( \sqrt{ \log d } ) Divergence

Authors: Guni Sharon, Alan Kuhnle

Organizations: Department of Computer Science and Engineering, Texas A&M University, College Station, TX 77843, USA.

Abstract

Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback--Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher--Rao path. We quantify this overhead by the inefficiency ratio R≥1R \ge 1, the Fisher length of the mixture geodesic divided by the Fisher--Rao distance, and bound its supremum over endpoint pairs as a function of the parameter dimension dd. A tensor criterion identifies the regime (I) families, with R=1R=1 everywhere: exactly those with quadratic potential or dimension one, such as fixed-covariance Gaussians. For non-quadratic families, we prove two further regimes: (II) bounded third-order skewness plus finite Fisher--Rao diameter yields a dimension-independent bound; and (III) for products of scale families---including Gaussian covariances and Gamma rates---RR grows as Θ(log⁡d)Θ(\sqrt{\log d}), unbounded in dd. Under a per-step Fisher-chord budget, RR translates to a practical computational cost: NGD requires asymptotically at least RR times as many steps as an optimizer following the Fisher--Rao geodesic. Experiments confirm all three regimes: R=1R=1 to machine precision for quadratic-potential families (I), the categorical bound π/(22)π/(2\sqrt{2}) is approached but not attained (II), and sampled scale-product RR grows with dd, reaching R≈1.5R \approx 1.5 for long, high-dimensional moves (III).

Explore similar work

CardsList
  1. The cost of useful natural gradient updates

    Sep 27, 2026Subhransu S. Bhattacharjee, Dylan Campbell, Rahul ShomeFisher Information MatrixGradient

  2. Natural gradient descent with momentum

    Apr 16, 2026Anthony Nouy, Agustín SomacalGradient DescentGradient

  3. Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients

    Mar 28, 2024Johannes Müller, Semih Çaycı, Guido MontúfarPolicy GradientFisher Information Matrix