math.NASep 4, 2026

Shallow neural network approximation in mixed Sobolev spaces

Authors: Yuwen Li, Guozhi Zhang

Organizations: School of Mathematical Sciences, Zhejiang University, 866 Yuhangtang Road, Hangzhou 310058, Zhejiang, China

Abstract

We investigate the best L2L_2 approximation of mixed Sobolev spaces by shallow neural networks with nn neurons and general activation functions. We first establish an activation-independent Fourier-block principle: if an activation has univariate approximation order ρρ in the sense of the Fourier-block property, then the global approximation rate has algebraic order min⁡{α,ρ}\min\{α,ρ\} for target functions of mixed smoothness αα, up to explicit logarithmic factors. To verify this property for concrete activations, we introduce a structured univariate approximation condition that implies the Fourier-block property with explicit parameters. For ReLUk\mathrm{ReLU}^k, a matching algebraic lower bound identifies min⁡{α,k+1}\min\{α,k+1\} as the optimal algebraic approximation exponent in any dimension, up to logarithmic factors in the upper bound. The framework also yields the exponent min⁡{α,k+1}\min\{α,k+1\} for cardinal B-splines and soft-ReLUk\mathrm{ReLU}^k, and the full mixed-smoothness exponent αα for ELU and cosine activations, again up to logarithmic~factors.

Figures & tables

Explore similar work

May 29, 2026stat.ML

Approximation and learning of anisotropic and mixed smooth functions by deep ReLU neural networks

This paper studies how efficiently deep ReLU neural networks can approximate and learn smooth functions. When the error is measured in Lp([0,1]d)L^p([0,1]^d) norm and the approximator is a network with width WW and depth LL, recent works have proven the supper approximation rate O((WL)−2s/d)\mathcal{O}((WL)^{-2s/d}) for Besov space Bq,rs([0,1]d)\mathcal{B}^s_{q,r}([0,1]^d) under the Sobolev embedding condition s/d>1/q−1/ps/d>1/q-1/p. In order to overcome the curse of dimensionality in this rate, we extent this result to anisotropic and mixed smooth function classes. We establish the approximation rate O((WL)−2s~)\mathcal{O}((WL)^{-2\tilde{s}}) for anisotropic Besov space Bq,rs([0,1]d)\mathcal{B}^{\boldsymbol{s}}_{q,r}([0,1]^d) with anisotropic smoothness s=(s1,…,sd)\boldsymbol{s}=(s_1,\dots,s_d) under the embedding condition s~>1/q−1/p\tilde{s} > 1/q-1/p, where the mean smoothness s~=(∑i=1dsi−1)−1\tilde{s} = (\sum_{i=1}^d s_i^{-1})^{-1}. For mixed smooth Besov space MBq,rs([0,1]d)\mathcal{MB}^s_{q,r}([0,1]^d) with mixed smoothness s>1/q−1/ps>1/q-1/p, we show that the approximation rate O((WL)−2s)\mathcal{O}((WL)^{-2s}) holds up to logarithmic factors. Using these results, we also derive approximation bounds for the composition of anisotropic Besov functions. As an application, it is shown that deep ReLU neural networks can achieve minimax optimal rates up to logarithmic factors for a wide range of smooth function classes.
Oct 5, 2025math.NA

Configuration-Dependent Lower Bounds for Approximation by Shallow ReLUk^k Networks on the Sphere

We establish two related but logically distinct results for shallow ReLUk^k neural networks on the unit sphere \SSd\SS^d. First, for an arbitrary set of inner neural-network parameters, the best L2(\SSd)\mathcal{L}^2(\SS^d) approximation of a fixed target function with smoothness r>d+2k+12r>\tfrac{d+2k+1}{2} admits an asymptotic lower bound given by a constant multiple of n−1/2h‾k+1/2n^{-1/2}\underline{h}^{k+1/2}, where h‾\underline{h} denotes the antipodal separation distance of the normalized inner-parameter set. This lower bound depends explicitly on the parameter configuration through h‾\underline{h} and applies without additional assumptions on the parameters. Second, for antipodally quasi-uniform parameters, h‾≃n−1/d\underline{h}\simeq n^{-1/d}, and the lower bound establishes the exact saturation order n−d+2k+12dn^{-\frac{d+2k+1}{2d}} for such parameter families: a target function with regularity greater than d+2k+12\frac{d+2k+1}{2} and satisfying the required parity condition can be approximated at this rate, whereas approximation at any strictly faster rate forces the target function to be zero. Our results therefore place linearized neural-network approximation within the classical saturation framework and show that, although ReLUk^k network spaces can outperform finite elements of the same degree, this advantage is intrinsically limited.
Jun 15, 2026stat.ML

Sobolev Approximation by Fixed-Size Neural Networks with Arbitrary Accuracy

In this work, we investigate new activation functions for achieving arbitrary-accuracy Sobolev approximation by fixed-size neural networks. We first show that any function in W2,∞((a,b)d)W^{2,\infty}((a,b)^d) can be approximated with arbitrary accuracy, measured in the W1,∞W^{1,\infty}-norm, by a fixed-size neural network using the Elementary Universal Activation Function (EUAF\mathrm{EUAF}). To extend this result to Ws,∞((a,b)d)W^{s,\infty}((a,b)^d) for s∈Ns\in\mathbb{N}, we introduce a smooth activation DUAF∞\mathrm{DUAF}_{\infty} from the family of Differentiable Universal Activation Functions (DUAFn\mathrm{DUAF}_n). We prove that any function in Ws,∞((a,b)d)W^{s,\infty}((a,b)^d) can be approximated with arbitrary accuracy in the Ws−1,∞W^{s-1,\infty}-norm by a fixed-size DUAF∞\mathrm{DUAF}_{\infty}-activated network. We further construct sigmoidal variants DUAF~n\widetilde{\mathrm{DUAF}}_n and show that, for every 1≤s≤n1\leq s\leq n, fixed-size DUAF~n\widetilde{\mathrm{DUAF}}_n-activated networks still approximate any f∈Ws,∞((a,b)d)f\in W^{s,\infty}((a,b)^d) with arbitrary accuracy in the Ws−1,∞W^{s-1,\infty}-norm. In all these results, the width and depth bounds are computed explicitly, and the proposed activations are elementary.