Sharp Approximation Rates for Neural Networks with Affine Latent Parameterizations
Organizations: Department of Applied Mathematics Hong Kong Polytechnic University
Abstract
Many parameter-efficient methods generate the parameters of a large neural network from a low-dimensional latent representation. Given an architecture with parameter slots, we write , where is a parameter generator and is a latent representation of the target function . The architecture and the generator are shared across the entire target class, while each target is represented by its own latent vector , with approximating . This framework encompasses hypernetworks, low-dimensional parameterizations, parameter-efficient adaptation, and model compression. Understanding the tradeoff between the latent dimension and the network budget is therefore fundamental to characterizing the expressive efficiency of these methods. We study this tradeoff for affine generators and fully connected ReLU architectures. More precisely, optimizing jointly over architectures satisfying and affine generators , we prove that the optimal worst-case uniform approximation error over the unit ball of -Hölder functions on , where , has the sharp order In particular, our result shows that even a fixed-dimensional latent space suffices to achieve vanishing approximation error as the network budget increases.