Function-Space Transformer with Adaptive Anchors
Organizations: University of Illinois Chicago
Abstract
Many forms of data, including physical fields, geometric shapes, and visual signals, are naturally described by functions over continuous domains but are observed through discrete samples. Representing these functions on fixed uniform grids imposes a trade-off between resolving localized variation and increasing computation across the domain. Neural operators address this mismatch by learning mappings between functions, while latent-attention architectures provide flexible processing of sampled observations. We introduce the Function-Space Transformer (FST), a framework for learning from functions through a spatially adaptive continuous latent representation. FST stores features at anchors whose locations are predicted from the input observations and recursively refines these anchor features through function-space interactions. This allows the representation to adapt its spatial organization to each input rather than inherit that of the observation grid, while supporting both spatially resolved and finite-dimensional outputs. On PDE solution prediction using PDEBench Burgers and Darcy flow, FST substantially outperforms the Perceiver IO baseline, whose latent representation lacks explicit spatial organization, and is highly competitive with the Fourier Neural Operator. On ImageNet-1K, FST achieves higher classification accuracy than the Vision Transformer baseline, with fewer parameters across these comparisons. Ablations further support the benefits of function-space updates and recursive refinement. Together, these results highlight the potential of adaptive continuous representations for both scientific prediction and visual recognition.
Figures & tables
| Model | Latent representation | Global interaction |
|---|---|---|
| Perceiver IO / UPT | Unstructured latent tokens | Self-attention among latent tokens |
| ENF | Latent features optimized for each input with Gaussian windows | Message passing among latents (ENF-PDE) |
| GPO | Predicted Gaussian geometry | Attention among modal tokens |
| AB-UPT | Spatial anchor tokens; separate query tokens | Self-attention among anchors; queries only read anchors |
| FST (ours) | Anchors predicted from the input, normalized RBFs | Attention among sampled field evaluations, written back |
| Model | Supervision per step | Parameters | Stop step | Relative |
|---|---|---|---|---|
| FST | 2,500 points | 2.901M | 500,000 | 0.018180 |
| FNO | full grid | 2.978M | 400,000 | 0.018528 |
| Perceiver IO | 2,500 points | 2.995M | 900,000 | 0.073293 |
| Model | Parameters | Supervision per step | ||||
|---|---|---|---|---|---|---|
| FST | 2.831M | 2,500 points | 0.04118 | 0.04159 | 0.04171 | 0.04185 |
| FNO | 2.978M | full grid | 0.02854 | 0.02950 | 0.02963 | 0.02974 |
| Perceiver IO | 2.937M | 2,500 points | 0.25634 | 0.25465 | 0.25396 | 0.25381 |
| Transolver | 2.956M | 2,500 points | 0.17709 | 0.07151 | 0.06569 | 0.06154 |
| DeepONet | 2.998M | 2,500 points | 0.24712 | 0.25204 | 0.25467 | 0.25581 |
| Coordinates | Length scales | anchors | anchors |
|---|---|---|---|
| Uniform | Fixed | 0.04604 | 0.04571 |
| Adaptive | Fixed | 0.04069 | 0.05784 |
| Uniform | Adaptive | 0.05482 | 0.05094 |
| Adaptive | Adaptive | 0.04185 | 0.04480 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Burgers | Darcy | ImageNet-1K |
| Anchors | |||
| Anchor feature dimension | 176 | 176 | 200 |
| Main attention heads | 8 | 8 | 8 |
| Observation-attention bias | Zero | Zero | Eq. 11 |
| Refinement steps | 4 | 4 | 4 |
| Evaluation points | 512 | 512 | 512 |
| Component | Burgers | Darcy | ImageNet-1K |
|---|---|---|---|
| Coordinate predictor | Fourier coordinate embeddings | Fourier coordinate embeddings | Fourier embeddings of patch-center coordinates |
| Predictor observation inputs | Coordinates, values, viscosity | Coordinates, coefficient values | Patch centers, patch features |
| Setting | All viscosities | |
|---|---|---|
| Training samples | 24,000 | 2,000 |
| Validation samples | 1,200 | 100 |
| Test samples | 1,000 | 84 |
| Original grid (time space) | ||
| IC values | 1,024 | |
| BC values | ||
| Model | Input | Supervision per training step |
|---|---|---|
| FST | ICs and BCs + | 2,500 sampled points |
| Perceiver IO | ICs and BCs + | 2,500 sampled points |
| FNO | Values + mask + on the grid | Full grid (205,824 points) |
| Model | Width | Layers | Heads | Latents/slices | Modes |
|---|---|---|---|---|---|
| FNO | 48 | 4 | - | - | |
| Perceiver IO | 240 | 6 | 8 | 256 | - |
| Transolver | 328 | 5 | 8 | 32 | - |
| Experiment | Anchors | Coordinates | Length scales | |
|---|---|---|---|---|
| All viscosities | Adaptive | Adaptive | 4 | |
| Single viscosity | Adaptive | Adaptive | 4 | |
| Default ablation | Adaptive | 0.03 | 4 |
| Setting | Value |
|---|---|
| Optimizer | Adam |
| Global batch size | 48 |
| Forward / backward precision | bfloat16 (BF16) |
| Loss accumulation precision | 32-bit floating point (FP32) |
| Minimum training steps | 250,000 |
| Maximum training steps | 1,000,000 |
| Model | Selected step | Stop step |
|---|---|---|
| FST | 92,000 | 98,000 |
| FNO | 72,000 | 72,000 |
| Perceiver IO | 34,000 | 50,000 |
| Transolver | 80,000 | 82,000 |
| DeepONet | 32,000 | 48,000 |
| Coordinates | Scales | ||
|---|---|---|---|
| Uniform | Fixed | 82k | 116k |
| Adaptive | Fixed | 86k | 80k |
| Uniform | Adaptive | 130k | 112k |
| Adaptive | Adaptive | 92k | 78k |
| Model | GPUs | Time (h) | Memory (GiB) |
|---|---|---|---|
| FST | 1 | 30.5 | 8.80 |
| FNO | 3 | 32.9 | 15.21 |
| Perceiver IO | 3 | 18.6 | 2.32 |
| FST | FNO | Perceiver IO | |
|---|---|---|---|
| 0.001 | 0.044195 | 0.048623 | 0.147162 |
| 0.002 | 0.033692 | 0.036266 | 0.120887 |
| 0.004 | 0.027352 | 0.031828 | 0.112725 |
| 0.01 | 0.021308 | 0.024882 | 0.088623 |
| 0.02 | 0.013957 | 0.015056 | 0.063128 |
| 0.04 | 0.013382 | 0.013013 | 0.061419 |
| Split | Samples | Grid | Precision |
|---|---|---|---|
| Training | 2,000 | BF16 | |
| Validation | 100 | FP32 | |
| Test | 1,000 | FP32 |
| Setting | Value |
|---|---|
| Training / validation images | 1,281,167 / 50,000 |
| Training crop | Random resized, |
| Training augmentations | Horizontal flips, 3-Augment ( Touvron et al., 2022 ) |
| Color jitter | 0.3 |
| Random erasing probability | 0.25 |
| Validation resize (short side) | 256 |
| Stage | Operation | Setting |
|---|---|---|
| 1 | Convolution | 64 channels |
| 2 | Residual blocks | 2 blocks, 2 convolutions each |
| 3 | Pointwise projection | Width 200 |
| 4 | Average pooling | Within each patch |
| 5 | Add positional embedding | Fourier embeddings of patch bounds + MLP |
| 6 | Residual feed-forward network | - |
| Setting | Value |
|---|---|
| Patch projection | Linear, patches |
| Feature width | 176 |
| Transformer layers | 10 |
| Attention heads | 4 |
| Feed-forward hidden width | 423 |
| Trainable parameters | 3,095,614 |
| Setting | Value |
|---|---|
| Optimizer | AdamW |
| Weight decay | 0.05 |
| Effective batch size | 256 |
| Computation precision | BF16 |
| Gradient norm clipping | 1.0 |
| Label smoothing | 0.1 |
| Stage | Start | End | Learning rate |
|---|---|---|---|
| Warmup | 0 | 25,023 | Peak |
| Cosine decay | 25,023 | 500,456 | |
| Cosine tail | 500,456 | 550,456 | |
| Constant learning rate | 550,456 | 600,456 |