Near-optimal estimates for the -Lipschitz constants of deep random ReLU neural networks
Abstract
This paper studies the -Lipschitz constants of ReLU neural networks with random parameters for . The distribution of the weights follows a variant of the He initialization. In the case of zero-bias networks, we derive high probability upper and lower bounds for wide networks that differ at most by a factor that is logarithmic in the network's depth. Remarkably, the behavior of the -Lipschitz constant varies significantly between the regimes and . For , the -Lipschitz constant behaves similarly to , where is a -dimensional standard Gaussian vector and . In contrast, for , the -Lipschitz constant aligns more closely to . We extend our analysis to networks with possibly non-zero biases drawn from arbitrary symmetric distributions. In this case, we obtain high probability upper and lower bounds that differ at most by a factor that is logarithmic in the network's width and linear in its depth.