Neural fields are usually evaluated by how well they reconstruct an observation. We show that this misses two useful properties of a fitted network: how easily it can adapt to new observations, and what its weights retain from earlier ones. We study these properties as adaptation geometry. For images, we meta-learn class-specific initializations, adapt each one to a new image, and measure how much the network must change to fit it. A simple local linear model closely predicts this adaptation cost, while replacing one network's tangent kernel with another's substantially worsens the prediction. Adaptation thus depends on the local geometry of the fitted network, not only on its current reconstruction. For physical fields, we repeatedly fit the same network to observations from a sequence. Its weights then retain information about that history. When two wave histories end at exactly the same observation, the final weights recover the sign of the wave velocity with 68.6% accuracy, whereas the current observation alone contains no such information and gives 50%. These two phenomena are quantitatively linked: tangent-kernel eigenvalues predict both which changes are easy to learn and how quickly they are overwritten by later fitting. Together, these results show that neural fields contain useful information beyond what they currently reconstruct: in how they can change and in how they got there.
Figures & tables
Figure 1: Adaptation geometry and sequential fitting carry information beyond reconstruction. (a) A class prior is its reconstruction μc and tangent kernel Kc that sets which output changes are cheap; reassignment keeps μc and swaps Kc . (b) In all settings, the matched kernel predicts nonlinear adaptation far better than a reassigned one (Spearman correlation with nonlinear adaptation energy). (c) Independent encoding restarts every frame from the shared prior θˉ ; sequential encoding starts each fit from the previous frame’s weights. (d) With the same class-conditioned readout, sequential weights classify the regime better than raw fields or independent weights on Allen–Cahn, Gray–Scott, and shear-flow (seed means); raw frames with explicit delays match them on Allen–Cahn but not on Gray–Scott. The shear flow delay search selects current frames alone (Table 2 ).
Spearman ρ
Decision agreement (%)
Setting
Δ Adapt.
Shared
Match
Shuffle
Shared
Match
Shuffle
MNIST-30 M-layer
+12.2
0.829
0.939
0.690
73.3
84.4
61.9
MNIST-100 M-layer
+5.3
0.886
0.955
0.770
75.7
88.3
58.9
Fashion-100 M-layer
+2.0
0.918
0.906
0.791
77.0
89.3
61.9
KMNIST-100 M-layer
+0.0
0.884
0.935
0.771
65.0
79.7
47.7
MNIST-30 SIREN
+4.4
0.618
0.873
0.327
38.9
57.8
11.0
Table 1: The correct local geometry predicts nonlinear adaptation. Correlation measures similarity in score rankings; agreement measures selection of the same prior. Shared uses the mean kernel; Match uses each prior’s own kernel; Shuffle reassigns kernels while fixing initial reconstructions (100 derangements). Δ Adapt. is nonlinear-minus-static classification accuracy in percentage points, a separate measure of utility. Entries are pooled held-out metrics rather than seed means; paired uncertainty for the method differences is reported in Appendix B.4 .
Figure 2: Different histories stay distinguishable at the same observed endpoint. (a) Paired wave histories differ in their earlier frames but end at exactly the same displacement field with opposite hidden velocities. (b) Velocity-sign accuracy of a linear probe on sequential endpoint weights as a function of how many frames were fitted before the endpoint; the observed field and independently fitted weights sit at chance by construction, and the field decoded from the sequential weights and its residual reach 55.6% and 56.4%. Dots at 16 frames are the nine seed cells; the whisker is the paired-bootstrap 95% interval.
System
Raw fields
Independent weights
Sequential weights
Raw + delays
Seq. minus best
Allen–Cahn
42.1 / 48.8
31.9 / 37.5
64.4 / 66.2
65.4
−1.0
Gray–Scott
57.0 / 53.3
59.0 / 55.6
64.9 / 65.6
59.3
+5.7
Shear flow (45 support/class)
27.7
28.6 / 33.0
34.8 / 44.6
27.7 ( H=1 )
+6.2
Table 2: Sequential fitting makes weight sequences more useful for classifying dynamics. Accuracy (%) as seed mean / score ensemble; an ensemble standardizes each query’s class-score vector within a run and averages across runs before deciding. Raw + delays selects among stacks of up to 16 raw frames in total: the current frame and at most 15 preceding frames. It reports means over three data seeds for Allen–Cahn and Gray–Scott, and one deterministic run for shear flow, where validation selects H=1 (no added history). The delay baseline uses the primary evaluation, whose Gray–Scott raw comparator is 56.3 rather than the diagnostic 57.0 shown here. The final column is the sequential seed mean minus the best other mean or deterministic result; differences are computed before rounding. Chance is 25% for Allen–Cahn and shear flow and 16.7% for Gray–Scott. Allen–Cahn and Gray–Scott cross three data seeds with three encoder seeds; shear flow uses one support split and three encoder seeds.
System
Steps/frame
Seq. MSE
Seq. accuracy
Allen–Cahn
50
0.059
64.6
200
0.043
69.4
1000
0.035
60.1
Gray–Scott
50
0.679
69.2
100
0.599
68.8
1000
0.395
64.1
Table 3: More fitting improves reconstruction but eventually hurts the history-bearing representation. Validation results for sequential fitting. Reconstruction error decreases monotonically with the number of fitting steps, while regime-classification accuracy peaks at an intermediate budget.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Five-field M-layer
Intensity-only SIREN
Outputs
5×28×28=3920
28×28=784
Adapted parameters
913
2,241
Prior initialization
Common within each fold
Common within each fold
Reptile inner steps
8 Adam steps
8 steps
Reptile inner learning rate
0.04
5×10−4
Reptile outer steps
1,000
1,000
Appendix
Table 4: Architecture and fitting settings for the image experiments.
Setting
Replication unit
Comparisons
Δρ
95% CI
Δ Agree
95% CI
MNIST-30 M-layer
Query bootstrap
900 pairs
+0.110
[0.093, 0.130]
+11.1
[3.3, 18.9]
Fold bootstrap
900 pairs
+0.110
[0.090, 0.132]
+11.1
[2.2, 18.9]
MNIST-100 M-layer
Query bootstrap
3,000 pairs
+0.069
[0.061, 0.078]
+12.7
[8.7, 16.7]
Fold bootstrap
3,000 pairs
+0.069
[0.060, 0.078]
+12.7
[8.3, 17.0]
MNIST-100 M-layer
27 prior sets
2,700 pairs
+0.069
[0.048, 0.095]
+13.7
[5.6, 23.7]
MNIST-100 SIREN
27 prior sets
2,700 pairs
+0.134
[0.096, 0.192]
+33.0
[22.6, 43.3]
Appendix
Table 5: Replication of the own-kernel advantage over shared geometry. Δρ is the correlation difference; Δ Agree is the decision-agreement difference in percentage points. Intervals are 95% confidence intervals. The last two rows refit the priors.
Spearman ρ
Decision agreement
Setting
Match
Shuffle
Match
Shuffle
MNIST-30 M-layer
0.943
0.698
82.2
62.6
MNIST-100 M-layer
0.957
0.770
89.0
58.5
Fashion-100 M-layer
0.909
0.797
87.3
62.3
KMNIST-100 M-layer
0.935
0.778
79.0
48.5
MNIST-30 SIREN
0.841
0.362
64.4
10.9
Appendix
Table 6: The reassignment effect persists after equalizing total kernel scale. Each kernel is rescaled to a common trace. Shuffle averages 100 derangements; decision agreement is in percent.
Architecture
Own
Same-class donor
Different-class donor
Shared
M-layer
0.9117
0.8539
0.5549
0.7385
SIREN
0.7847
0.7142
−0.0802
0.0917
Appendix
Table 9
Representation
Test (ensemble)
Encoder seeds
Macro-F1
Stored params
Raw states / selected delay ( H=1 )
27.7
n/a
0.233
4,224
Independent INR weights
33.0
33.9 / 23.2 / 28.6
0.235
4,612
Sequential endpoints ( β=10−3 )
44.6
33.0 / 38.4 / 33.0
0.439
676
Sequential endpoints ( β=0 )
45.5
33.9 / 38.4 / 33.9
0.439
676
Full adaptation path (6 checkpoints)
39.3
33.0 / 37.5 / 31.2
0.384
2,116
Decoded sequential fields
26.8
23.2 / 24.1 / 31.2
0.225
n/a
Appendix
Table 7: Shear flow results and representation controls. Accuracy (%) is shown for the score ensemble and individual encoder seeds. Raw states and the validation-selected delay baseline are the same deterministic representation ( H=1 ). Stored parameter counts include the dynamics readout for one encoder seed, excluding PCA arrays.
Representation
Mean
Worst seed
Seed SD
Range
Ensemble
Raw fields
56.3
54.4
1.39
3.3
57.8
Sequential weights
64.9
62.2
1.18
3.3
65.6
Appendix
Table 8: Gray–Scott held-out stability. Accuracy (%) and variation across encoder/data seeds for the primary held-out evaluation.
Data seed
Enc. seed
Sequential
Seq. decoded
Residual
Velocity R2
Pair RMSE
14011
101
67.2
56.6
58.2
0.229
0.0062
14011
211
68.0
54.7
50.8
0.216
0.0066
14011
307
68.0
52.0
58.6
0.217
0.0065
25031
101
64.8
57.4
59.0
0.234
0.0064
25031
211
67.6
57.0
54.7
0.222
0.0066
25031
307
68.4
52.7
53.5
0.200
0.0065
Appendix
Table 9: Wave-memory results for each seed cell. Sign accuracy is in percent. Current-field and independent-weight probes are exactly 50% in every row. Velocity R2 uses sequential weights; Pair RMSE measures the difference between paired sequential reconstructions.
Representation / update
Seed mean
Ensemble
Seed SD
Raw-state dynamics
57.0
53.3
3.8
Independent nonlinear INR
59.0
55.6
4.0
Sequential nonlinear INR
64.9
65.6
1.2
Matched tangent update
46.5
54.4
4.1
Mismatched tangent update
30.1
42.2
6.3
Appendix
Table 10: Replacing the Jacobian changes a local update model on Gray–Scott. Accuracy (%) over nine seed cells, using 128 Jacobian points per transition. Matched beats mismatched in all nine cells. This diagnostic run reports its own raw comparator.
System
Instances
Median Spearman
Median log slope
Slow ret. s=16
Fast ret. s=16
Allen–Cahn
216
0.986
1.005
0.998
0.259
Gray–Scott
324
0.945
0.992
0.999
0.579
Appendix
Table 11: Predicted and measured retention under controlled updates. Correlation compares mode-wise retention for each transition. A log-retention slope of one is exact agreement. Slow and fast columns give retention after 16 steps.
System
Steps/frame
Seq. MSE
Indep. acc.
Seq. acc.
Allen–Cahn
1
0.695
26.7
25.7
Allen–Cahn
50
0.059
37.2
64.6
Allen–Cahn
200
0.043
42.7
69.4
Allen–Cahn
1000
0.035
45.1
60.1
Gray–Scott
1
1.128
52.2
56.6
Gray–Scott
50
0.679
55.5
69.2
Appendix
Table 12: More fitting improves reconstruction but not always classification. Validation results. Raw accuracy is fixed at 41.7% for Allen–Cahn and 61.1% for Gray–Scott.
Diagnostic
Ideal
Result
Median complete-field relative closure (fine RK4)
0
2.99×10−14
Closure reduction, coarse → fine integration
large
16.5×
Net parameter-change local slope in loop scale
2
2.023
Tangent-kernel-change local slope
2
1.979
Forward/reverse parameter cosine
−1
−0.999954
Direct Lie-bracket cosine
+1
0.9948
Appendix
Table 13: Full controlled-loop diagnostics. Local medians and slopes use nine base-field/initialization combinations at the two smallest loop scales, 0.005 and 0.01 (18 combination/scale records); figure-eight suppression uses the available control records at these scales. The minimum measured tangent-kernel change covers all nine combinations and all four scales.
Department of Computer Science, University of Haifa · Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology · School of Computer Science and AI, Tel Aviv University