ReLU Neural Networks

ReLU: Rectified Linear Unit

Momentum

12 papers in the last four weeks, up 300% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 63

Jun 24, 2025stat.ML

Near-optimal estimates for the ℓp\ell^p-Lipschitz constants of deep random ReLU neural networks

This paper studies the ℓp\ell^p-Lipschitz constants of ReLU neural networks Φ:Rd→RΦ: \mathbb{R}^d \to \mathbb{R} with random parameters for p∈[1,∞]p \in [1,\infty]. The distribution of the weights follows a variant of the He initialization. In the case of zero-bias networks, we derive high probability upper and lower bounds for wide networks that differ at most by a factor that is logarithmic in the network's depth. Remarkably, the behavior of the ℓp\ell^p-Lipschitz constant varies significantly between the regimes p∈[1,2)p \in [1,2) and p∈[2,∞]p \in [2,\infty]. For p∈[2,∞]p \in [2,\infty], the ℓp\ell^p-Lipschitz constant behaves similarly to ∥g∥p′\Vert g\Vert_{p'}, where g∈Rdg \in \mathbb{R}^d is a dd-dimensional standard Gaussian vector and 1/p+1/p′=11/p + 1/p' = 1. In contrast, for p∈[1,2)p \in [1,2), the ℓp\ell^p-Lipschitz constant aligns more closely to ∥g∥2\Vert g \Vert_{2}. We extend our analysis to networks with possibly non-zero biases drawn from arbitrary symmetric distributions. In this case, we obtain high probability upper and lower bounds that differ at most by a factor that is logarithmic in the network's width and linear in its depth.
Feb 4, 2025stat.ML

Networks with Finite VC Dimension: Pro and Contra

Approximation and learning of classifiers of large data sets by neural networks in terms of high-dimensional geometry and statistical learning theory are investigated. The influence of the VC dimension of sets of input-output functions of networks on approximation capabilities is compared with its influence on consistency in learning from samples of data. It is shown that, whereas finite VC dimension is desirable for uniform convergence of empirical errors, it may not be desirable for approximation of functions drawn from a probability distribution modeling the likelihood that they occur in a given type of application. Based on the concentration-of-measure properties of high dimensional geometry, it is proven that both errors in approximation and empirical errors behave almost deterministically for networks implementing sets of input-output functions with finite VC dimensions in processing large data sets. Practical limitations of the universal approximation property, the trade-offs between the accuracy of approximation and consistency in learning from data, and the influence of depth of networks with ReLU units on their accuracy and consistency are discussed.
Jul 13, 2023cs.LG

Deep Network Approximation: Beyond ReLU to Diverse Activation Functions

This paper explores the expressive power of deep neural networks for a diverse range of activation functions. An activation function set A\mathscr{A} is defined to encompass the majority of commonly used activation functions, such as ReLU\mathtt{ReLU}, LeakyReLU\mathtt{LeakyReLU}, ReLU2\mathtt{ReLU}^2, ELU\mathtt{ELU}, CELU\mathtt{CELU}, SELU\mathtt{SELU}, Softplus\mathtt{Softplus}, GELU\mathtt{GELU}, SiLU\mathtt{SiLU}, Swish\mathtt{Swish}, Mish\mathtt{Mish}, Sigmoid\mathtt{Sigmoid}, Tanh\mathtt{Tanh}, Arctan\mathtt{Arctan}, Softsign\mathtt{Softsign}, dSiLU\mathtt{dSiLU}, and SRS\mathtt{SRS}. We demonstrate that for any activation function ϱ∈A\varrho\in \mathscr{A}, a ReLU\mathtt{ReLU} network of width NN and depth LL can be approximated to arbitrary precision by a ϱ\varrho-activated network of width 3N3N and depth 2L2L on any bounded set. This finding enables the extension of most approximation results achieved with ReLU\mathtt{ReLU} networks to a wide variety of other activation functions, albeit with slightly increased constants. Significantly, we establish that the (width, \,depth) scaling factors can be further reduced from (3,2)(3,2) to (1,1)(1,1) if ϱ\varrho falls within a specific subset of A\mathscr{A}. This subset includes activation functions such as ELU\mathtt{ELU}, CELU\mathtt{CELU}, SELU\mathtt{SELU}, Softplus\mathtt{Softplus}, GELU\mathtt{GELU}, SiLU\mathtt{SiLU}, Swish\mathtt{Swish}, and Mish\mathtt{Mish}.