Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred from observed data. In meta-learning, this encoding occurs through gradient updates induced by the decoder, tying representation learning directly to decoder design. We formalize this connection by interpreting latent optimization as an optimization encoder, unifying the roles of second-order differentiation, latent parameterization, and task supervision. This concept enables second-order meta-learning for end-to-end training of the encoding procedure alongside the decoder, and clarifies which learning pathway first-order approximations discard. Guided by this view, we introduce Attentive Latent Fields (MetaLF), an equivariant transformer-based neural field that contextualizes a latent pointcloud through self-attention. These interactions shape both field predictions and the updates that construct their representation, allowing local observations to inform coherent non-local structure. Disentangling the inner encoding objective from outer task supervision unifies reconstruction, classification, and segmentation within an end-to-end meta-learning framework, using reconstruction-only latent adaptation at test time. Controlled experiments on polynomial fields link latent coordination to lower effective rank and stronger alignment with the underlying function space. Across image and 3D shape reconstruction, MetaLF improves fidelity within three to five gradient updates, while supporting semantic prediction across images, shapes, and volumes. Together, these findings position the optimization encoder perspective as a unified basis for designing neural fields around how representations are constructed, coordinated, and used.
Figures & tables
Figure 1: Example use case of MetaLF meta-learning latent cardiac shape representations [ 3 ] . The first three panels illustrate smooth shape recovery as a shared latent pointcloud z is refined over k=1,…,5 gradient steps. MetaLF learns contextualized latent representations that preserve detail while capturing the underlying function, an ill-posed task for local representation models such as ENF [ 4 ] .
Figure 2: Second-order meta-learning trains an optimization encoder whose behaviour is shaped by decoder architecture. (a) A conventional encoder maps an observed signal to a latent representation zi , which the decoder evaluates at a query coordinate x . (b) In meta-learned neural fields, an explicit encoder is replaced by K gradient updates on z , forming an optimization encoder. Second-order differentiation propagates the outer objective Lout through these unrolled updates, jointly shaping the finite-step encoding behavior and the decoder fθ . (c) Because the decoder induces the latent updates, its architecture shapes the resulting adaptation trajectory. ENF represents a local fitting model, lacking direct exchange of context among localized latents. MetaLF instead coordinates these latents through attention. This shared context produces coherent updates that quickly capture non-local structure, as illustrated by fitting cardiac signed distance functions.
Figure 3: MetaLF as an optimization encoder. (a) K gradient updates encode observations into a spatially grounded, globally contextualized latent representation through SE(n) -equivariant self-attention (SA) layers and cross-attention (CA) with a query coordinate. (b) During meta-training, task-specific outer objectives can shape the decoder through second-order gradients across the adaptation steps. Global contextualization of the latent point set therefore enables flexible meta-training: the same inner-loop encoder can learn representations tailored to diverse downstream objectives.
Figure 4: A receiving latent j attends to neighbouring source latents j′ using bi-invariant geometric attributes aj′→j . Keys and values combine this geometry with the source content cj′ , and their attention-weighted aggregation produces the updated feature hj′ .
Full-grid PSNR ↑ (dB)
MAML geometry
MAML
FOMAML
reff↓
τpoly↑
Functa
37.31 (0.31)
35.03 (1.32)
128.27 (2.11)
0.539 (0.006)
ENF
44.15 (0.56)
32.95 (0.65)
118.14 (13.54)
0.418 (0.031)
MetaLF ( NSA=0 )
42.28 (0.13)
33.25 (0.62)
72.85 (4.78)
0.619 (0.035)
MetaLF ( NSA=2 )
63.79 (1.35)
37.40 (1.87)
33.31 (0.29)
0.837 (0.004)
MetaLF ( NSA=4 )
64.24 (0.82)
38.30 (0.76)
31.09 (0.31)
0.861 (0.004)
Table 1: MetaLF self-attention improves both encoding structure and reconstruction quality. Increasing NSA yields lower-dimensional, more polynomial-aligned encodings (effective rank reff , tangent fraction τpoly ) and higher decoded-field PSNR. Results are averaged over three seeds.
CIFAR10
CelebA
ImageNet
Functa
36.2 (0.1)
32.6 (0.1)
23.1 (0.0)
ENF
aR2
42.5 (0.3)
35.7 (0.1)
27.7 (0.0)
ENF
aSE(2)
42.4 (0.2)
35.7 (0.1)
27.6 (0.0)
MetaLF
aR2
47.7 (0.2)
39.4 (0.0)
32.0 (0.0)
MetaLF
aSE(2)
47.9 (0.1)
39.5 (0.1)
31.9 (0.0)
Table 2: Test-set reconstruction PSNR (dB, ↑ ) for CIFAR10, CelebA (64x64) and ImageNet-1K (128x128), averaged over three seeds.
ShapeNet IoU ↑
ACDC IoU ↑
Part
Core
LV
Myo
RV
Functa
4.0 (0.1)
5.0 (0.0)
1.2 (0.0)
1.1 (0.0)
1.2 (0.0)
SpatialFuncta
49.0 (2.2)
46.2 (2.3)
69.4 (5.5)
47.2 (2.6)
55.5 (2.0)
ENF
65.5 (0.5)
67.2 (1.9)
77.1 (2.1)
37.3 (2.6)
64.7 (1.1)
MetaLF
76.7 (1.5)
78.2 (4.0)
89.3 (0.5)
77.9 (0.4)
79.5 (2.7)
Table 3: MAML-based shape reconstruction IoU for the ShapeNet-Part, -Core, and ACDC hold-out test-sets. Values are the mean (standard deviation) over three seeds.
ImageNet
CIFAR10
ShapeNet-Core
PSNR ↑
Acc. ↑
PSNR ↑
Acc. ↑
IoU ↑
Acc. ↑
MWT
21.8 (-)
24.1 (-)
30.9 (-)
64.7 (-)
-
-
Functa
21.6 (2.5)
7.4 (0.1)
23.1 (1.9)
57.7 (0.5)
25.6 (8.0)
1.8 (9.4)
SpatialFuncta
25.5 (2.5)
4.1 (0.1)
36.5 (2.9)
42.6 (0.5)
48.6 (12.9)
64.8 (30.8)
ENF
26.3 (2.7)
4.5 (0.1)
39.2 (3.3)
34.4 (0.5)
48.8 (13.2)
57.3 (34.3)
MetaLF
27.9 (2.9)
42.6 (0.2)
37.8 (2.4)
83.3 (0.4)
55.3 (16.9)
70.8 (24.7)
Table 4: MAML-based classification and reconstruction on the CIFAR10, ImageNet-1K, and ShapeNet-Core hold-out test sets. PSNR (dB) and IoU are mean (standard deviation) over test items; image accuracy is top-1 (binomial standard error) and ShapeNet accuracy is macro-averaged over its 55 classes.
OMBRIA
ShapeNet-Part
OASIS
( K=3 )
( K=5 )
( K=5 )
IoU ↑
mIoU ↑
Dice ↑
OmbriaNet
72.4
–
–
MetaSeg ( K=100 )
–
–
91.0 (1.1)
NISF
–
–
81.0 (0.7)
Functa
39.9
Failed
Failed
Table 5: Meta-learned segmentation on the OMBRIA, ShapeNet-Part, and OASIS hold-out test sets. Header K denotes test-time fitting steps for our methods. Scores are the pooled flood IoU for OMBRIA, instance mIoU for ShapeNet-Part, and five-class foreground Dice for OASIS. Parentheses denote standard deviations over test items. Failed denotes non-fit/divergence under the attempted recipes.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: MetaLF is equivariant under a group action defined by its bi-invariant feature a . Rows show translation-equivariant ( Rn ) and roto-translation-equivariant ( SE(n) ) MetaLF models on CelebA and ACDC myocardium. Columns show ground truth, reconstruction in the canonical frame, pose-only translation, joint translation of poses and queries, pose-only roto-translation, and the corresponding joint roto-translation. Acting on poses alone transforms the represented field; applying the same group element jointly to poses and queries preserves predicted values at corresponding coordinates when the model respects that symmetry. Accordingly, joint translations match the canonical reconstruction in every row, whereas joint roto-translation does so only for the SE(n) models.
Figure 6: Training progress on CIFAR10 reconstruction for MAML (blue) and FOMAML (red). (a) MAML develops a higher latent effective rank, while the first-order controls remain at lower ranks. (b) MAML generally has smaller upper response factors, although these fluctuate during training. (c) MAML’s reconstruction error decreases substantially over training epochs, whereas the first-order controls plateau after an initial improvement. (d) Both methods perform best near the three-step training horizon, with reconstruction error gradually increasing under further adaptation.
Figure 7: Optimization encoder-side coupling (top) and decoder-side attention routing (bottom) evaluated on the polynomial fields dataset, shown with increasing number of MetaLF self-attention layers ( NSA ) for N=25 latent tokens. Additional layers allow information to propagate beyond local latent neighborhoods while retaining the spatial structure of the representation.
CIFAR10
CelebA
ImageNet-1K
Latent tokens N
25
36
169
Initial latent grid
5×5
6×6
13×13
(KSA,KCA)
(5,4)
(6,4)
(13,4)
Query/value frequency scales
(1,3)
(1,3)
(2,10)
Inner / outer coordinates
1,024/1,024
4,096/4,096
4,096/16,384
Effective batch size
32
32
64
Appendix
Table 6: MetaLF training settings for image reconstruction. Settings are shared by the aR2 and aSE(2) variants. The inner coordinate count is per adaptation step.
ShapeNet
ACDC
Latent tokens N
27
256
Initial latent grid
3×3×3
8×8×4
(KSA,KCA)
(3,8)
(6,8)
Near-surface fraction
75%
50%
Inner / outer coordinates
1,638/8,192
24,000/24,000
Training epochs
100
200
Appendix
Table 7: Shape-reconstruction settings. Coordinate counts refer to the SDF data term; ACDC additionally uses surface anchors. ShapeNet settings are shared by the 16-category and 55-category experiments.
Table 9: MetaLF meta-segmentation settings. Coordinate counts describe one training signal; ShapeNet-Part additionally supervises annotated surface points. The same number of inner steps is used during training and testing.
Figure 8: Qualitative example for OMBRIA meta-segmentation, illustrating the 25th percentile, median, and 75th percentile MetaLF test-set performance sample in terms of IoU. In the median case, SpatialFuncta and ENF produce extensive false-positive flood predictions around cloudy regions, whereas MetaLF largely suppresses them. This pattern is consistent with non-local latent interactions using broader spatial context to resolve ambiguous local signals.
Figure 9: Reconstruction quality as a function of latent pointcloud size. The example uses MetaLF with NSA=4 and per-latent content size of D=64 .
D=16
D=32
D=64
N
NSA=0
NSA=4
NSA=0
NSA=4
NSA=0
NSA=4
1
18.92
19.06
20.57
20.62
22.36
22.49
4
22.69
23.03
24.77
25.29
27.58
28.42
9
25.59
26.47
28.61
29.77
32.72
34.20
16
28.47
29.85
32.54
34.17
39.17
40.78
25
31.29
33.18
36.67
38.82
46.44
48.28
Appendix
Table 10: CIFAR-10 validation set reconstruction PSNR for MetaLF capacity variations. Rows vary the number of latents N ; grouped columns vary latent channels D and self-attention depth NSA .
Reconstruction ( B=1 )
Meta-training ( B=32 )
Model
Params (M) ↓
GFLOPs ↓
ms/image ↓
GFLOPs ↓
ms/update ↓
Functa
6.187
11.975
1.383 [1.382, 1.390]
1147.983
40.367 [40.364, 40.419]
Spatial Functa
0.354
3.846
0.572 [0.570, 0.573]
368.079
14.543 [14.537, 14.557]
ENF
0.610
33.232
1.870 [1.868, 1.870]
3188.983
116.081 [116.032, 116.099]
MetaLF ( NSA=0 )
0.362
10.748
1.219 [1.218, 1.220]
1031.587
50.863 [50.836, 50.993]
MetaLF ( NSA=2 )
1.120
11.435
3.173 [3.170, 3.174]
1088.456
58.189 [58.082, 58.206]
Appendix
Table 11: Compute on CIFAR10 ( 32×32 ) with three full-image inner loop steps and 1,600 content scalars, with ENF and MetaLF additionally fitting 50 pose scalars. Spatial Functa uses 5×5×64 latents. Parameters include the shared decoder, trainable initialization, and trainable adaptation rates. Reconstruction includes encoding and decoding, while training includes a full second-order meta-update. Timings use one NVIDIA H100 GPU and report the median with the interquartile interval below, across 100 synchronized calls. FLOPs count logical XLA arithmetic (multiply-add =2 ), excluding transcendentals. Bold and underline denote the lowest and second-lowest costs.
Figure 10: Uncurated examples generated using latent rectified flow based on the reconstruction meta-learned MetaLF (MetaLF-Diffuser) representation models.
CIFAR10
CelebA
Model
FID ↓
KID ×103↓
Cov. (%) ↑
FID ↓
KID ×103↓
Cov. (%) ↑
ENF (DiT)
18.58
6.89 ±0.14
93.06
18.82
17.82±0.18
70.51
MetaLF (DiT)
20.08
7.22±0.18
90.56
13.05
11.05±0.12
86.43
MetaLF (MetaLF-Diffuser)
18.83
6.59±0.19
92.87
13.38
11.57 ±0.14
86.53
Appendix
Table 12: Generative modeling on CIFAR10 and CelebA using 50,000 generated samples, evaluated against 10,000 and 19,962 test images, respectively. CIFAR10 models are class-conditioned; CelebA models are unconditional. KID is reported as mean ± block standard error.
Neural fields parameterize data as functions from coordinates to values, providing a unified framework for representation learning across modalities. Existing approaches are dominated by per-sample meta-learning, which scales poorly due to memory-intensive inner-loop optimization. The natural alternative -- feed-forward encoding -- typically introduces modality-specific assumptions, sacrificing the generality that makes learning with neural fields attractive. We argue that locality and hierarchy are useful priors for learning field representations that can be injected without compromising modality-agnosticism. We propose LH-NeF, a framework to learn general-purpose tokenized representations of continuous signals. A locality-preserving hierarchical encoder maps raw coordinate-value field observations to structured tokens, from which the field is reconstructed during training. By replacing meta-learning's inner loop with a single forward pass, LH-NeF uses 42× less memory and supports 133× larger batches than the strongest modality-agnostic baseline. Across images, 3D shapes, and climate fields, our learned representations match or exceed performance of modality-agnostic, modality-specific, and specialized generative neural field baselines on both reconstruction and downstream tasks.
Alonso Urbano, David W. Romero, Max Zimmer +1
Department for AI in Society, Science, and Technology, Zuse Institute Berlin (ZIB), Germany · Cartesia AI, San Francisco, CA, USA · Institute of Mathematics, Technische Universität Berlin, Germany
Neural fields (NFs) map continuous coordinates to signals such as color or density, but fast high-quality reconstruction from sparse observations remains difficult. Classical Neural Tangent Kernel (NTK) regression gives closed-form fits, yet it is fundamentally linear and cannot accumulate reusable task priors. We develop three algorithms that address these gaps. NTK-KIP learns a distilled support set of coordinates (and optional labels) so that a finite NTK can inpaint large missing regions from little observed data, yielding a compact non-linear representation instead of a raw kernel solve. MetaQuill meta-learns a shared initialization for an INR so that new scenes can be adapted by updating only a small task-specific weight offset, which provides true feature learning and a reusable prior. Finally, MetaQuill-KIP fuses both ideas: it seeds the task with a KIP-style non-linear warm start, then refines only that small offset around the meta-learned initialization. MetaQuill-KIP achieves high-PSNR reconstructions and semantically plausible inpainting under very sparse observations, while requiring only lightweight per-instance adaptation, whereas diffusion-style baselines typically depend on large pretrained generative priors and costly per-image tuning. This shows that NTK-driven neural fields can be made both non-linear and meta-learnable, narrowing the gap between analytic kernels and practical few-shot reconstruction.
Amir Mallak, Alaa Maalouf, Lior Wolf +2
Department of Computer Science, University of Haifa · Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology · School of Computer Science and AI, Tel Aviv University
We investigate the potential of weights to serve as effective representations, focusing on neural fields. Our key insight is that constraining the optimization space through a pre-trained base model and low-rank adaptation (LoRA) can induce structure in weight space. Across reconstruction, generation, and analysis tasks on 2D and 3D data, we find that multiplicative LoRA weights achieve high representation quality while exhibiting distinctiveness and semantic structure. When used with latent diffusion models, multiplicative LoRA weights enable higher-quality generation than existing weight-space methods.