math.OC · 2607.20411 Copy arXiv ID · Jul 22, 2026 Save Lipschitzian SLLNs for random functions Authors: Lai Tian , Johannes O. Royset
Organizations: Daniel J. Epstein Department of Industrial and Systems Engineering, University of Southern California, Los Angeles, CA.
Abstract We prove strong laws of large numbers for locally Lipschitz functions in the Lipschitz pseudometric. Our results hold under either a topological or a model-theoretic condition, with the latter encompassing functions jointly definable in o-minimal structures but extending substantially beyond this class. Applications include uniform convergence of limiting and Clarke subdifferentials and finite-sample identification of solutions. Consequently, we identify broad classes of functions for which the failure phenomena revealed by our previous negative results [Tian and Royset, arXiv:2511.16568, 2025] do not occur.
Explore similar work Sep 2, 2026 · Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios Lipschitz Constant Metric Spaces
May 28, 2026 · Marius Potfer, Vianney Perchet Lipschitz Constant Multi-Armed Bandits
May 29, 2026 · Arnak S. Dalalyan, Avetik Karagulyan Langevin Dynamics Non-Log-Concave
Sep 2, 2026 · stat.ML J/K move · Enter open · S save
Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios
Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks. We address this by introducing a simple closed-form ``two-stage'' compositional formula
f ^ \hat{f} f ^ for reconstructing an unknown Lipschitz function
f : X → R f:\mathcal{X}\to \mathbb{R} f : X → R on a metric space
( X , ρ ) (\mathcal X,ρ) ( X , ρ ) from
N N N i.i.d. noisy observations. Our main result is a high-probability uniform (
L ∞ L^{\infty} L ∞ ) recovery guarantee that jointly controls approximation and statistical errors while enjoying an optimization error of zero; in particular, we do not assume oracle access to an approximate ERM. Our secondary main results establish the optimality of our formula in three complementary senses. 1) Function space: On Ahlfors-regular metric spaces, the hypothesis class parameterized by our formula attains the optimal fat-shattering dimension. 2) Parameter space: Its dependence on the parameters is maximally numerically stable, in the sense that a smaller approximation error cannot be achieved with a smaller Lipschitz dependence on the model parameters. 3) Forward pass: Its dependence on the input is maximally regular, matching the Lipschitz constant of the target function
f f f . When
X = [ 0 , 1 ] d \mathcal X=[0,1]^d X = [ 0 , 1 ] d is equipped with the
ℓ ∞ \ell^\infty ℓ ∞ norm,
f ^ \hat{f} f ^ admits algorithmic ReLU-MLP and exact ReLU-multi-head transformer realizations of depth
O ( log ( N ) ) \mathcal{O}(\log(N)) O ( log ( N )) with
O ( N ) \mathcal{O}(N) O ( N ) nonzero parameters.