Organizations: School of Mathematics, Georgia Institute of Technology, Atlanta, GA 30332, USA · Department of Mathematics, Emory University, Atlanta, GA 30322, USA · School of Computational Science and Engineering, Georgia Institute of Technology, Atlanta, GA 30332, USA
Transformers have emerged as powerful architectures for learning solution operators of physical systems. Empirically the prediction error has been observed to decrease when the data size and model size increase, suggesting neural scaling behavior. Yet a theoretical understanding of such scaling laws for transformer-based operator learning remains limited. In this work, we develop a theoretical framework for characterizing the approximation and generalization errors of transformer-based operator learning. Our analysis builds on a local-to-global approximation principle that is naturally aligned with the softmax attention mechanism and yields discretization-invariant output functions. On approximation theory, we derive a universal approximation error of transformer-based operator learning for Hölder-regular operators. On generalization theory, we establish a power scaling law between the prediction error and the training data size. The rate of convergence represented by the scaling exponent explicitly reflects the dimensions of the input and output domains, the regularity of the underlying functions and operators, and crucially, the intrinsic dimension of the input function class. By exploiting this intrinsic low-dimensional structure, our analysis yields a power-law generalization rate for operator learning, in contrast to the logarithmic-type power-law rates appearing in existing analyses of operator learning with feedforward neural networks. Numerical experiments validate the predicted power-law scaling and confirm that the convergence rate varies systematically with the intrinsic dimension of the input function class.
Figures & tables
Figure 1 : Illustration of the operator-learning data-generation setting. Each input function is observed on a shared input grid, while output values are queried at independently sampled locations.
Figure 2: Illustration of the two-level Softmax POU approximation. The input-space POU localizes u over function-space anchors {uk}k=1CU , while the output-domain POU localizes the query y over anchors {zl}l=1CV . Their combination yields the oracle approximation in ( 23 ).
Figure 3: Empirical data scaling for Burgers’ (left) and KdV (right). Markers show the geometric mean test MSE over five repeats, with error bars corresponding to one sample standard deviation of the log MSE; dashed lines show the fitted power law MSE≈Cn−p .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Burgers’
KdV
Initial learning rate
10−3
3×10−4
Batch size
256
64
Random output queries per training function
8
16
Maximum epochs
12,000
6,000
Test functions
128
128
Test spatial grid
1024 uniform points
1024 uniform points
Appendix
Table 1: Training and evaluation settings for the Burgers’ and KdV experiments.
Figure 4: Representative Burgers’ predictions. The top and bottom rows correspond to n=32 and n=1024 , respectively; the left and right columns correspond to dU=2 and 8 . Dotted curves show inputs, solid curves reference solutions, and dashed curves model predictions. Each column uses the same held-out test input across both sample sizes.
Figure 5: Representative KdV predictions. The top and bottom rows correspond to n=32 and n=1024 , respectively; the left and right columns correspond to dU=4 and 8 . Dotted curves show inputs, solid curves exact solutions, and dashed curves model predictions. Each column uses the same held-out test input across both sample sizes.
We study the approximation of nonlinear operators between function spaces by transformers. Our approach is to lift functions to measures supported on their graphs and leverage a recently introduced measure-theoretic view of transformers. A function h is represented by its graph measure γh, with finite tokens {(xj,h(xj))}j=1N being its empirical approximations. We show that this framework elegantly models discretization refinement via convergence of measures and provides a natural setting for operator learning. Within this framework, we introduce function graph transformers, a graph-preserving subclass of measure-theoretic transformers that maps graph measures to graph measures, which is to say that outputs remain single-valued functions. Crucially, this additional structure does not reduce generality: we prove that the resulting graph-preserving maps can be approximated by finite compositions of standard softmax self-attention layers and pointwise MLPs, yielding universal approximation results for broad classes of nonlinear operators. Unlike existing theoretical approaches to operator learning with transformers, the measure-theoretic framework also accommodates regularized negative-order Sobolev inputs for which discretization invariance is particularly challenging, as well as query points on different output domains. Overall, function graph transformers provide a continuum viewpoint and mathematical toolkit for transformer-based operator learning, clarifying the roles of positional encodings, graph structure, regularization, and ensuring consistency across discretizations.
Takashi Furuya, David Mis, Ivan Dokmanić +2
Doshisha University, RIKEN AIP · Rice University · University of Basel +2
We develop approximation and generalization error estimates for multi-input neural operators, with the output error measured in Sobolev norms. In contrast to standard operator-learning settings with a single input function, our framework allows multiple input functions defined on possibly different domains, with different dimensions and Sobolev regularities. The derived rates explicitly quantify the contribution of each input space to the final error bound. In particular, in the balanced regime, the approximation and generalization rates are governed by the interaction between the input dimensions, regularities, and Sobolev orders, while the dependence on the model complexity retains a loglog/log-type structure. Our analysis provides a general theoretical framework for multi-input operator learning, including Sobolev training, and is applicable to operator learning problems arising from partial differential equations and scientific computing.
Yahong Yang, Zecheng Zhang, Wei Zhu +2
School of Mathematics, Georgia Institute of Technology, 686 Cherry Street, Atlanta, GA 30332-0160, USA. · Department of Applied and Computational Mathematics and Statistics, University of Notre Dame, Notre Dame, IN 46556, USA. · Department of Mathematics, Hong Kong Baptist University, FSC1202, Fong Shu Chuen Building, Hong Kong Baptist University, Kowloon Tong, Hong Kong.
This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain [0,1]d and d-dimensional compact Riemannian manifolds. We propose a novel constructive approximation framework for Transformers that builds local approximations of the target function and aggregates them into a global approximation via softmax partition of unity. This approach leverages the attention mechanism to achieve spatial localization through affine transformations of the input. The softmax activation plays a crucial role in aggregating local approximations to a global output. From an approximation perspective, we prove that a dense Transformer equipped with only two encoder blocks and standard single-hidden-layer point-wise feed-forward networks can achieve a uniform ε-approximation error for α-Hölder continuous functions with α∈(0,1] using O(ε−d/α) total parameters. Building upon this approximation guarantee, we establish a near minimax-optimal generalization error bound of order O(n−2α+d2αlogn) for the empirical risk minimizer, where n is the training data size. The Transformer architecture studied in this paper is dense, shallow and wide, and employs softmax activation and sinusoidal positional encodings, closely reflecting practical implementations.
Zhongjie Shi, Wenjing Liao
School of Mathematics, Georgia Institute of Technology, Atlanta, GA 30332, United States