Many forms of data, including physical fields, geometric shapes, and visual signals, are naturally described by functions over continuous domains but are observed through discrete samples. Representing these functions on fixed uniform grids imposes a trade-off between resolving localized variation and increasing computation across the domain. Neural operators address this mismatch by learning mappings between functions, while latent-attention architectures provide flexible processing of sampled observations. We introduce the Function-Space Transformer (FST), a framework for learning from functions through a spatially adaptive continuous latent representation. FST stores features at anchors whose locations are predicted from the input observations and recursively refines these anchor features through function-space interactions. This allows the representation to adapt its spatial organization to each input rather than inherit that of the observation grid, while supporting both spatially resolved and finite-dimensional outputs. On PDE solution prediction using PDEBench Burgers and Darcy flow, FST substantially outperforms the Perceiver IO baseline, whose latent representation lacks explicit spatial organization, and is highly competitive with the Fourier Neural Operator. On ImageNet-1K, FST achieves higher classification accuracy than the Vision Transformer baseline, with fewer parameters across these comparisons. Ablations further support the benefits of function-space updates and recursive refinement. Together, these results highlight the potential of adaptive continuous representations for both scientific prediction and visual recognition.
Figures & tables
Model
Latent representation
Global interaction
Perceiver IO / UPT
Unstructured latent tokens
Self-attention among latent tokens
ENF
Latent features optimized for each input with Gaussian windows
Message passing among latents (ENF-PDE)
GPO
Predicted Gaussian geometry
Attention among modal tokens
AB-UPT
Spatial anchor tokens; separate query tokens
Self-attention among anchors; queries only read anchors
FST (ours)
Anchors predicted from the input, normalized RBFs
Attention among sampled field evaluations, written back
Table 1: Where each architecture stores its latent representation and where it computes global interaction.
Figure 1: Function-Space Transformer (FST). (a) A shared layer refines anchor features for R steps at fixed coordinates XA , predicted once from the input. The decoder queries the field at Q or pools anchor features. (b) Each step attends to observations, evaluates RBF blends through a pointwise MLP at XE(r) , and applies self-attention. Cross-attention updates anchors using queries from the same RBF blend and MLP at XA .
Figure 2: Evaluate, mix, and write-back. (a) RBF weights blend anchor features; a pointwise MLP, omitted for clarity, maps the blend and coordinate embedding to an evaluated feature. (b) Evaluation points may include anchor coordinates. (c) Self-attention mixes evaluated features. (d) Cross-attention incorporates information into anchor features. Its queries use the same RBF blend and MLP at anchor coordinates. Anchor coordinates stay fixed; one receiving anchor is highlighted.
Model
Supervision per step
Parameters
Stop step
Relative L2
FST
2,500 points
2.901M
500,000
0.018180
FNO
full grid
2.978M
400,000
0.018528
Perceiver IO
2,500 points
2.995M
900,000
0.073293
Table 2: Burgers results across all 12 viscosities. Stop step: end of training. Selected steps (lowest validation error): FST 450k, FNO 400k, Perceiver IO 895k.
Figure 3: Per-viscosity test error on PDEBench Burgers (all-viscosity models). (a) Relative L2 error. (b) 100(eFST−eFNO)/eFNO ; negative values favor FST. All 12 viscosities are shown.
Figure 4: Burgers profiles at ν=0.001 (top) and ν=0.02 (bottom). Colors indicate time; each row shares axis limits. FST captures fronts with fewer spurious oscillations (top) but rounds initial-state minima (bottom).
Model
Parameters
Supervision per step
32×32
64×64
128×128
201×1024
FST
2.831M
2,500 points
0.04118
0.04159
0.04171
0.04185
FNO
2.978M
full grid
0.02854
0.02950
0.02963
0.02974
Perceiver IO
2.937M
2,500 points
0.25634
0.25465
0.25396
0.25381
Transolver
2.956M
2,500 points
0.17709
0.07151
0.06569
0.06154
DeepONet
2.998M
2,500 points
0.24712
0.25204
0.25467
0.25581
Table 3: Test relative L2 error on Burgers ( ν=0.01 ) at four output grids. See Appendix B for detailed inputs, outputs, and evaluation.
Coordinates
Length scales
322 anchors
162 anchors
Uniform
Fixed
0.04604
0.04571
Adaptive
Fixed
0.04069
0.05784
Uniform
Adaptive
0.05482
0.05094
Adaptive
Adaptive
0.04185
0.04480
Table 4: Ablation on anchor coordinates and length scales.
Table 9Table 10
Figure 5: Full-grid relative L2 error using different numbers of nearby anchors.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Converting predicted gaps to anchor coordinates. Each row has its own predicted spacing.
Setting
Burgers
Darcy
ImageNet-1K
Anchors M
32×32=1,024
32×32=1,024
32×32=1,024
Anchor feature dimension
176
176
200
Main attention heads
8
8
8
Observation-attention bias
Zero
Zero
Eq. 11
Refinement steps R
4
4
4
Evaluation points S
512
512
512
Appendix
Table 9: FST architecture settings for the reported experiments.
Component
Burgers
Darcy
ImageNet-1K
Coordinate predictor
Fourier coordinate embeddings
Fourier coordinate embeddings
Fourier embeddings of patch-center coordinates
Predictor observation inputs
Coordinates, values, viscosity
Coordinates, coefficient values
Patch centers, patch features
Appendix
Table 10: Coordinate predictor encodings and observation inputs.
Setting
All viscosities
ν=0.01
Training samples
24,000
2,000
Validation samples
1,200
100
Test samples
1,000
84
Original grid (time × space)
201×1024
IC values
1,024
BC values
2×200=400
Appendix
Table 11: Burgers datasets and shared observation grid.
Figure 7: Burgers decoding at one query coordinate.
Model
Input
Supervision per training step
FST
ICs and BCs + ν
2,500 sampled points
Perceiver IO
ICs and BCs + ν
2,500 sampled points
FNO
Values + mask + ν on the 201×1024 grid
Full grid (205,824 points)
Appendix
Table 12: Model inputs and output supervision for the Burgers comparison.
Table 14: FST configurations for the Burgers comparisons and the default ablation. The default settings are varied in Tables 4 - 6 .
Setting
Value
Optimizer
Adam
Global batch size
48
Forward / backward precision
bfloat16 (BF16)
Loss accumulation precision
32-bit floating point (FP32)
Minimum training steps
250,000
Maximum training steps
1,000,000
Appendix
Table 15: Shared training settings for the all-viscosity Burgers comparison.
Model
Selected step
Stop step
FST
92,000
98,000
FNO
72,000
72,000
Perceiver IO
34,000
50,000
Transolver
80,000
82,000
DeepONet
32,000
48,000
Appendix
Table 16: Checkpoints for the single-viscosity models in Table 3 .
Coordinates
Scales
322
162
Uniform
Fixed
82k
116k
Adaptive
Fixed
86k
80k
Uniform
Adaptive
130k
112k
Adaptive
Adaptive
92k
78k
Appendix
Table 17: Selected training steps for the Burgers ablations; k denotes 1,000 steps. Left: coordinate and length-scale ablations at two anchor counts. Right: update and refinement ablations with 322 anchors.
Model
GPUs
Time (h)
Memory (GiB)
FST
1
30.5
8.80
FNO
3
32.9
15.21
Perceiver IO
3
18.6
2.32
Appendix
Table 18: All-viscosity training on NVIDIA A100 80GB GPUs. Time sums the recorded training segments, including validation. Memory is the peak allocated memory per GPU.
ν
FST
FNO
Perceiver IO
0.001
0.044195
0.048623
0.147162
0.002
0.033692
0.036266
0.120887
0.004
0.027352
0.031828
0.112725
0.01
0.021308
0.024882
0.088623
0.02
0.013957
0.015056
0.063128
0.04
0.013382
0.013013
0.061419
Appendix
Table 19: Per-viscosity relative L2 error on PDEBench Burgers. Best in bold; second best underlined.
Figure 8: Additional Burgers solution profiles. From top to bottom, rows show selected trajectories at ν=0.001 , 0.002 , 0.004 , 0.01 , and 0.02 . Columns show the reference solution and predictions from FST, FNO, and Perceiver IO. Colors indicate t∈{0,0.1,0.3,0.6,1,2} ; axis limits are shared across columns within each row. The first and last rows are the examples in Figure 4 . Different rows have different initial conditions.
Figure 9: Burgers solution profiles for different initial conditions at fixed viscosity ν=0.002 . From top to bottom, the initial profiles are Gaussian-like, single wave, two waves, and localized oscillations. Columns show the reference solution and predictions from FST, FNO, and Perceiver IO. Colors indicate t∈{0,0.1,0.3,0.6,1,2} ; axis limits are shared across columns within each row. The examples contrast broad spatial features, repeated wave structures, and spatially concentrated oscillations.
Split
Samples
Grid
Precision
Training
2,000
128×128
BF16
Validation
100
100×100
FP32
Test
1,000
128×128
FP32
Appendix
Table 20: Darcy data and training settings.
Figure 10: Darcy anchor coordinates and decoding RBF length scales for four test samples. (a) Anchors over reference solutions. (b) Soft ellipses with horizontal and vertical scales proportional to ℓx and ℓy , displayed at 0.35× scale; color indicates ℓxℓy .
Setting
Value
Training / validation images
1,281,167 / 50,000
Training crop
Random resized, 224×224
Training augmentations
Horizontal flips, 3-Augment ( Touvron et al., 2022 )
Color jitter
0.3
Random erasing probability
0.25
Validation resize (short side)
256
Appendix
Table 21: ImageNet-1K and data preprocessing.
Stage
Operation
Setting
1
Convolution
64 channels
2
Residual blocks
2 blocks, 2 convolutions each
3
Pointwise projection
Width 200
4
Average pooling
Within each patch
5
Add positional embedding
Fourier embeddings of patch bounds + MLP
6
Residual feed-forward network
-
Appendix
Table 22: FST patch encoder, in execution order. Spatial convolutions use 3×3 kernels and the encoder uses GELU activations.
Figure 11: ImageNet patch observations. A CNN encodes each 16×16 patch. Spatial pooling, a patch-bound embedding, and a residual FFN produce the feature hi , paired with the patch center ci=((xi0+xi1)/2,(yi0+yi1)/2) .
Figure 12: Anchor coordinates and RBF length scales on ImageNet-1K validation images. (a, c) All 1,024 anchors overlaid on the images. (b, d) RBF length scales ℓref , shown by color and soft circles displayed at 0.35× scale. Horizontal coordinates adapt to each input; vertical rows are fixed.
Learning mappings between infinite-dimensional function spaces, or operator learning, is essential for many machine learning applications. Although transformer-based operators are popular, they often rely on token-wise attention. These methods treat continuous fields as discrete tokens and usually ignore the global functional structure. We introduce \emph{Functional Attention}, which reinterprets attention as a functional correspondence between adaptive bases. Inspired by geometric functional maps, our method replaces softmax affinities with structured linear operators. This yields a compact, generalizable, resolution-invariant representation that explicitly captures global dependencies. Experiments demonstrate that \emph{Functional Attention} can match state-of-the-art performance in many operator learning tasks, including solving PDEs, 3D segmentation, and regression, while remaining robust to varying discretizations. Project page is available at https://github.com/xjffff/FUNCATTN.
Jiefang Xiao, Maolin Gao, Simon Weber +2
Technical University of Munich, Germany · Munich Center for Machine Learning (MCML), Germany · PIXL, Department of Computer Science, University of Oxford, United Kingdom +1
We study the approximation of nonlinear operators between function spaces by transformers. Our approach is to lift functions to measures supported on their graphs and leverage a recently introduced measure-theoretic view of transformers. A function h is represented by its graph measure γh, with finite tokens {(xj,h(xj))}j=1N being its empirical approximations. We show that this framework elegantly models discretization refinement via convergence of measures and provides a natural setting for operator learning. Within this framework, we introduce function graph transformers, a graph-preserving subclass of measure-theoretic transformers that maps graph measures to graph measures, which is to say that outputs remain single-valued functions. Crucially, this additional structure does not reduce generality: we prove that the resulting graph-preserving maps can be approximated by finite compositions of standard softmax self-attention layers and pointwise MLPs, yielding universal approximation results for broad classes of nonlinear operators. Unlike existing theoretical approaches to operator learning with transformers, the measure-theoretic framework also accommodates regularized negative-order Sobolev inputs for which discretization invariance is particularly challenging, as well as query points on different output domains. Overall, function graph transformers provide a continuum viewpoint and mathematical toolkit for transformer-based operator learning, clarifying the roles of positional encodings, graph structure, regularization, and ensuring consistency across discretizations.
Takashi Furuya, David Mis, Ivan Dokmanić +2
Doshisha University, RIKEN AIP · Rice University · University of Basel +2
Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in latent spaces. However, we reveal that existing learnable projection mechanisms cannot ensure stable and balanced assignments from observation points to latent tokens, causing some latent tokens to be over-assigned while others remain underutilized. This limitation further restricts the design of hierarchical architectures, as assignment imbalance is continuously inherited and amplified across latent spaces, eventually causing severe token collapse in deeper spaces. To address these issues, we propose MoNo (Multiscale Optimal Transport Neural Operator), a progressive multiscale neural operator that efficiently solves PDEs on general geometries through stable latent-space construction. At its core is CoTAP (Cross-scale Optimal Transport Assignment and Projection), a novel latent-space construction method that formulates cross-space assignment between adjacent spaces as an entropy-regularized optimal transport problem, thereby constructing balanced bidirectional projections and stable latent spaces. CoTAP also ensures stable information transfer across multiple latent spaces, further enabling multiscale architectures on general geometries, which in turn support more efficient learning of long-range physical interactions. Extensive experiments demonstrate that MoNo outperforms existing state-of-the-art neural operators in both prediction performance and computational efficiency. Code is available at https://github.com/ZijiangY1116/MoNo.
Zijiang Yang, Xiaomeng Wu, Dongmei Fu
School of Automation and Electrical Engineering, University of Science and Technology Beijing · 2Beijing Engineering Research Center of Industrial Spectrum Imaging