cs.AIAug 15, 2026
SaveA concentration result for multilayer feedforward neural networks
Abstract
We consider for an arbitrary fixed and for each positive integer a multilayer feedforward artificial neural network with layers, neurons in the first layer (the input layer) and only one neuron, the output neuron, in the last layer. Very roughly formulated, the main result is that if the distribution of weights of connections from a layer to the next are, for all large , approximated well by a fixed continuous (but otherwise arbitrary) curve which does not depend on , and if the values of the input neurons are independently and identically distributed with a continuous probability density function, then there is a number such that for all the probability that the value of the output neuron is in tends to 1 as tends to infinity.
Explore similar work
Universal approximation theorems provide a mathematical explanation for the expressive power of neural networks. They assert that, under mild conditions on the activation function, feedforward neural networks are dense in broad function classes, such as continuous functions on compact subsets of , spaces, or Sobolev spaces. Over the past four decades, these qualitative universality results have evolved into a rich quantitative theory addressing approximation rates, parameter efficiency, and the role of architectural features such as depth and width. This survey presents several glimpses into this theory. We review classical density results for single-hidden-layer networks, as well as quantitative bounds that relate approximation error to network size and smoothness assumptions on target functions. Particular emphasis is placed on depth--width trade-offs and on results demonstrating that deeper architectures can achieve superior parameter efficiency for structured function classes. In addition to standard feedforward neural networks, we also review recent developments on Kolmogorov--Arnold Networks (KANs), which offer an alternative architectural paradigm and whose approximation-theoretic properties have begun to attract significant theoretical attention.
A law of robustness for two-layer neural networks with arbitrary weights
Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with neurons that fits noisy labels must have Lipschitz constant at least of order , with no restriction on the size of the weights. Bubeck and Sellke proved a universal version of this law for Lipschitz-parameterized classes, but under a polynomial bound on the parameters; at depth three that boundedness hypothesis is genuinely necessary. The two-layer unbounded-weight case requires a different argument. We prove the conjectured law, up to one logarithmic factor, for every continuous piecewise-linear activation, in particular for ReLU networks. For data drawn uniformly from , , or from , labels in with noise level , and any width- two-layer network with arbitrary real weights, biases and affine skip connection, fitting the data below the noise floor forces , , with high probability. A realized-kink-count version holds on the same event: every realized two-layer piecewise-linear function with distinct kink hyperplanes obeys the bound with replaced by , irrespective of how many redundant hidden units parameterize it. The proof replaces parameter-space covering, impossible for unbounded weights, by a function-space covering. The central deterministic ingredient is a rigidity lemma: on , and on for , the coefficient of each canonical kink is controlled by the Lipschitz constant of the realized function, because kinks on distinct hyperplanes cannot cancel at generic points. Rigidity genuinely fails at , and an explicit two-layer ReLU interpolant with Lipschitz constant at width matches the law at the overparameterized endpoint.
Optimal Non-Asymptotic Edgeworth Expansions for Multivariate Neural Network Outputs
Finite-width fully connected neural networks with Gaussian-initialized weights deviate from their infinite-width Gaussian limit, exhibiting non-vanishing higher-order cumulants. We approximate these deviations, for a neural network evaluated in a finite number of inputs, using multidimensional Edgeworth expansions of arbitrary order , with . Assuming that the corresponding Gaussian limit has an invertible covariance matrix and that the activation function is polynomially bounded, we establish a bound of order on the total variation distance between the law of the true network output and its Edgeworth approximation, with matching lower bounds. As an application, we quantify the error in Bayesian posterior distributions when the prior is replaced by its Edgeworth expansion. Our results are more general and also apply to sequences of conditionally Gaussian vectors converging to a Gaussian vector with invertible covariance.