cs.LGOct 1, 2026

Least-time Gradient Flow

Authors: Alessandro Betti, Marco Gori, Stefano Melacci, Jinwei Zhao

Organizations: Department of Mathematics “Tullio Levi-Civita”, University of Padua, Italy. · Department of Information Engineering and Mathematics, University of Siena, Italy. · University of Florence, Italy.

Abstract

Prescribing the speed of gradient flow on the risk itself, by the dynamics w˙=−u(E(w))∇E(w)/\abs∇E(w)2\dot w=-u(E(w))\nabla E(w)/\abs{\nabla E(w)}^{2}, makes the risk e(t)=E(w(t))e(t)=E(w(t)) obey e˙=−u(e)\dot e=-u(e) exactly, whatever the landscape~EE; the time needed to reach zero risk from e0e_0 is ∫0e0\dde/u(e)\int_0^{e_0}\dd e/u(e). Minimizing this time alone is ill posed, and we study the regularized problem inf⁡{∫0e0(λ2\absu′2+1/u) \dde: u∈H1(0,e0), u≥0, u(0)=0}\inf\{\int_0^{e_0}(\tfrac\lambda2\abs{u'}^{2}+1/u)\,\dd e:\ u\in H^{1}(0,e_0),\ u\ge0,\ u(0)=0\}, λ>0λ>0. We prove that the minimizer exists, is unique, and is a linearly scaled cycloid, and we show that the optimal rate behaves like u∗(e)∼(9/(2λ))1/3e2/3u^{*}(e)\sim(9/(2λ))^{1/3}e^{2/3} near zero risk: the exponent 2/32/3 is the one found in \cite{betti2026holder} by a power-law ansatz, and it lies in the Hölder window (12,1)(\tfrac12,1) where the arrival is in finite time with vanishing weight speed. The proof follows the classical route: existence by the direct method, uniqueness by strict convexity, positivity of the minimizer away from the origin, and the explicit integration of the Euler-Lagrange equation.

Explore similar work

CardsList
  1. The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow

    Sep 1, 2026Raphaël BerthierGradient DescentGradient

  2. Accelerating Min-Max Optimization via Power-Law Stepsizes

    Jun 1, 2026Yue Wu, Weiqiang Zheng, Yang Cai +1Step AccuracyConvergence