cs.LGSep 23, 2026

Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections

Authors: Karolina DrabikBen LewisAntoni PuchEtienne BoursierPiotr HofmanMatthias EnglertRanko Lazić

Abstract

We study minimal-norm interpolation and 2\ell_2-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-layer biases are included in the parameter norm. When biases are unpenalized, the minimal-norm interpolators are exactly the continuous piecewise-affine functions that hug every label switch and have kinks of the appropriate convexity. When biases are penalized, the minimizer is unique in function space, has exactly one kink in each intermediate same-label segment, and is therefore a sparsest positive-margin classifier. We further show that adding a free affine skip connection leaves these function-space solutions unchanged but fundamentally improves the parameter-space landscape: every KKT point of the constrained problem becomes globally optimal, whereas suboptimal KKT points can occur without the skip connection. We establish analogous global-optimality and geometric results for sufficiently weak 2\ell_2-regularization of the logistic loss. In the unpenalized-bias case, we identify an additional sparsity-like restriction, implying that most minimal-norm interpolators cannot arise as small-regularization limits of margin-normalized logistic-loss minimizers. Numerical experiments across varying dataset complexity and network width support the predicted landscape and sparsity phenomena.

Explore similar work

Jul 8, 2026cs.LG

A law of robustness for two-layer neural networks with arbitrary weights

Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with mm neurons that fits nn noisy labels must have Lipschitz constant at least of order n/m\sqrt{n/m}, with no restriction on the size of the weights. Bubeck and Sellke proved a universal version of this law for Lipschitz-parameterized classes, but under a polynomial bound on the parameters; at depth three that boundedness hypothesis is genuinely necessary. The two-layer unbounded-weight case requires a different argument. We prove the conjectured law, up to one logarithmic factor, for every continuous piecewise-linear activation, in particular for ReLU networks. For data drawn uniformly from Sd1\mathbb{S}^{d-1}, d3d\ge3, or from N(0,Id/d)N(0,I_d/d), labels in [1,1][-1,1] with noise level σ2>0σ^2>0, and any width-mm two-layer network with arbitrary real weights, biases and affine skip connection, fitting the data ε\varepsilon below the noise floor forces Lip(f)cεn/(mˉlog(Cmˉnd/ε))\mathrm{Lip}(f)\ge c\,\varepsilon\sqrt{n/(\bar m\log(C\bar m nd/\varepsilon))}, mˉ=(K1)m+1\bar m=(K-1)m+1, with high probability. A realized-kink-count version holds on the same event: every realized two-layer piecewise-linear function with k(f)nk(f)\le n distinct kink hyperplanes obeys the bound with mˉ\bar m replaced by k(f)+1k(f)+1, irrespective of how many redundant hidden units parameterize it. The proof replaces parameter-space covering, impossible for unbounded weights, by a function-space covering. The central deterministic ingredient is a rigidity lemma: on B2B_2, and on Sd1\mathbb{S}^{d-1} for d3d\ge3, the coefficient of each canonical kink is controlled by the Lipschitz constant of the realized function, because kinks on distinct hyperplanes cannot cancel at generic points. Rigidity genuinely fails at d=2d=2, and an explicit two-layer ReLU interpolant with O(1)O(1) Lipschitz constant at width 2n2n matches the law at the overparameterized endpoint.
Yitzchak Shmalo
May 26, 2026cs.LG

Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias

The successful training of neural networks hinges on the use of first order optimization methods, yet the theoretical characterization of these methods remains incomplete. This is especially true in settings with mild overparameterization. In this work, we study the gradient flow dynamics of two-layer ReLU networks from small initialization with orthogonal training data. We prove the limiting flow converges to a saddle-to-saddle jump process as the initialization scale tends to zero, revealing an incremental learning phenomenon in which a new neuron activates at each saddle. This analysis recovers the known result of Dana et al. (2025, arXiv:2502.16977) that the network interpolates the training data with high probability as soon as mlog(n)m \gtrsim \log(n), where mm is the network width and nn is the number of training samples. This incremental process characterization also allows us to derive a novel implicit bias result: the learned interpolator has a squared 2\ell_2-norm scaling as n\sqrt{n}, which is within a constant factor of the minimal 2\ell_2-norm interpolator. More broadly, our work provides the first rigorous proof of an incremental learning process for ReLU networks, whilst suggesting mildly overparameterized networks can converge to interpolating solutions whose complexity is of the same order as that of the optimal interpolator.
James Town, Etienne Boursier, Ben Lewis +2
Jul 18, 2026cs.LG

Effects of width-dependent model hyperparameters and 2\ell_2-regularization on the loss landscape of two-layer ReLU networks

Understanding deep neural networks remains a central challenge in machine learning. In particular, the theoretical properties of even two-layer ReLU networks, especially in the presence of weight decay, remain poorly understood. To this end, we derive a sufficient condition on the hyperparameter settings under which the global minima collapse to the zero solution. Interestingly, our experiments reveal that using AdamW as an optimizer prevents the collapse of the learned parameters, whereas using SGD does not, which may help explain the success of AdamW in deep learning training. In addition, when restricting the input dimension to one, we derive an analytical solution for the globally optimal parameter sets of two-layer ReLU networks and show that 2\ell_2-regularization has a width-invariant effect on connectivity, but its dimensionality-reducing effect becomes stronger as the network width increases. These results provide insight into how width-dependent hyperparameters influence the geometry of regularized loss landscapes.
Haruka Eshima, Makoto Yamada