math.OCSep 27, 2026

Local LMO is Secretly a Projection Method!

Authors: Peter Richtárik, Ammar Mahran

Organizations: King Abdullah University of Science and Technology, Thuwal, Saudi Arabia

Abstract

The local linear minimization oracle (Ferris and Zavriev, 1996; arXiv:2605.08850), or Local LMO, solves constrained convex problems without having to compute a projection: it minimizes a linear model over the intersection of the feasible set with a ball around the current iterate. We show that, whenever the ball radius does not exceed the Polyak radius, the Local LMO step is the Euclidean projection of the current iterate onto the intersection of the feasible set with a half-space that separates the iterate from the solution set; this projection is nonetheless computable by a linear oracle alone. We demonstrate that Local LMO belongs to a broader family of projection methods which may be indexed by the depth of the localizing half-space. For objectives with ϑ\vartheta-Hölder continuous gradient, every method from this family whose half-space lies sufficiently deep drives the best of its first KK iterates to optimality at the universal rate O(K−(1+ϑ)/2)\mathcal{O}(K^{-(1+\vartheta)/2}), matching the non-accelerated universal gradient method of Nesterov (2015). Run at the Polyak radius, Local LMO attains the same rate when the constrained optima are also unconstrained (∥∇f(x⋆)∥=0\|\nabla f(x_\star)\| = 0), and the rate O(K−1/(2−ϑ))\mathcal{O}(K^{-1/(2-\vartheta)}) otherwise.

Explore similar work

May 9, 2026math.OC

Broximal Gradient Descent: A Projection-Free Sister of Projected Gradient Descent

We propose Broximal Gradient Descent (BroxGD), a projection-free sister method to projected gradient descent for constrained optimization. Its forward--backward construction replaces the proximal backward operation by the broximal operation of Gruntkowska et al. (2025). Instead of projecting, each step minimizes a linear function over the intersection of the constraint set X\mathcal{X} and a ball B(xk,tk)\mathbb{B}(x_k,t_k) centered at the current iterate xkx_k, of suitable radius tk>0t_k>0: xk+1∈arg⁡min⁡z∈X∩B(xk,tk)⟨∇f(xk),z⟩.x_{k+1}\in\arg\min_{z\in\mathcal{X}\cap\mathbb{B}(x_k,t_k)}\langle\nabla f(x_k),z\rangle. We develop a comprehensive convergence theory spanning a wide range of optimization regimes and radius rules. We expect BroxGD to find many applications and inspire numerous extensions, much like projected gradient descent. Our contribution is theoretical; potential applications and toy experiments illustrate the method and suggest directions for future work.
Apr 18, 2026math.OC

Trajectory-Restricted Optimization Conditions and Geometry-Aware Linear Convergence

Linear convergence of first-order methods is typically characterized by global optimization conditions whose constants reflect worst-case geometry of the ambient space. In high-dimensional or structured problems, these global constants can be arbitrarily conservative and fail to capture the geometry actually encountered by optimization trajectories. In this paper, we develop a trajectory-restricted framework for linear convergence based on localized geometric regularity. We introduce restricted variants of the Polyak--Łojasiewicz inequality, error bound, and quadratic growth conditions that are required to hold only on subsets of the domain. We show that classical convergence guarantees extend under these localized conditions, and in key cases, we develop new arguments that yield explicit relationships between the corresponding constants. The resulting rates are governed by geometric quantities associated with the regions traversed by the algorithm. For polyhedral composite problems, we prove that convergence is controlled by restricted Hoffman constants corresponding to the active polyhedral faces visited along the trajectory. Once the iterates enter a well-conditioned face, the effective condition number improves accordingly. Our work provides a geometric quantification for fast local convergence after active-set or manifold identification and more broadly suggests that linear convergence is fundamentally governed by the geometry of the subsets explored by the algorithm, rather than by worst-case global conditioning.
Sep 16, 2026math.OC

Matching Multi-Loop Complexities with a Single Loop: Optimal Optimization Stationarity and Best-Known Game Stationarity in Nonconvex--Concave Minimax Optimization

We introduce a new single-loop algorithmic framework for smooth nonconvex--concave minimax optimization. The resulting projected damped extragradient method combines projected extragradient updates, dual momentum, and a moving proximal center. Under both the optimization-stationarity and game-stationarity criteria, our method achieves the best-known complexity among single-loop first-order methods. For optimization stationarity, our method achieves a gradient complexity of O(L2DYΔˉ0ε−3)O(L^2D_Y\barΔ_0\varepsilon^{-3}), where LL is the gradient Lipschitz constant, DYD_Y bounds the diameter of the dual feasible set, and Δˉ0\barΔ_0 is an initialization quantity involving the value-function gap and the initial gradients. Moreover, by incorporating a fixed-center warm-up phase, the complexity can be improved to O(L2DYΔφε−3)O(L^2D_YΔ_φ\varepsilon^{-3}), up to an additive lower-order cost, where Δφ:=φ(x0)−inf⁡xφ(x)Δ_φ:=φ(x_0)-\inf_xφ(x). We further establish a lower bound of Ω(L2DYΔφε−3)Ω(L^2D_YΔ_φ\varepsilon^{-3}) for optimization stationarity over projected zero-respecting first-order methods. This lower bound proves that the warm-started version of our algorithm is optimal up to a constant factor for optimization stationarity within this oracle class. For game stationarity, our method achieves O ⁣(L3/2DY1/2Δφε−5/2)\mathcal{O}\!(L^{3/2}D_Y^{1/2}Δ_φ\varepsilon^{-5/2}) gradient complexity. This matches the best-known complexity of multi-loop first-order methods, thereby establishing the same complexity with a single-loop algorithmic structure. Under dual strong concavity, the proposed framework achieves O ⁣(κ LΔφε−2)O\!(\sqrtκ\,LΔ_φ\varepsilon^{-2}) leading complexity for both stationarity criteria, where κ=L/μκ=L/μ is the dual condition number, up to an additive initialization cost. The ε−2\varepsilon^{-2} accuracy dependence is optimal under fixed regularity and initialization bounds.