Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression
Abstract
The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data. While increasing network size is a natural way to improve approximation, parameter magnitude provides another resource whose role must be quantified in both respects. We establish a sharp width--magnitude tradeoff at fixed depth using one elementary bounded -Lipschitz Dyadic--Triangular Activation. For the unit -Hölder ball on with , the optimal approximation error for is of order when the network width satisfies and the parameter magnitudes are bounded by . Matching lower bounds hold for every fixed globally Hölder activation; its Hölder exponent affects the constants but not the rate. Under bounded design densities and independent centered sub-Gaussian noise, approximate least squares over the full clipped class at depth attains the classical Hölder minimax risk without logarithmic loss whenever , where is the sample size. This yields a continuum of statistically optimal choices, ranging from unit parameter radius to fixed network size. At fixed size, four hidden layers with at most nonzero parameters give a near-optimal radius, while six layers with at most attain the optimal order at approximation error . The same decoding method also yields fixed-size Transformer approximation.