Edge Accuracy Is Not Enough: Why Dynamics-Learned Structure Fails to Transfer to Inverse Problems
Authors: Nicholas Tan Jerome, Fangnian Wang
Organizations: Institute for Data Processing and Electronics (IPE) Karlsruhe Institute of Technology (KIT) Karlsruhe, Germany · Institute for Thermal Energy Technology and Safety (ITES) Karlsruhe Institute of Technology (KIT) Karlsruhe, Germany
A natural strategy for inverse problems with scarce labelled data is to transfer relational structure learned from abundant forward-simulation data. We show this strategy fails systematically, even when it satisfies the standard theoretical justification for why structure should help. We prove that approximate structure provides estimation-error benefits whenever the edge error satisfies Δ<n2−kn, reducing sample complexity from O(n2) to O(kn+Δ). Structure learned via Neural Relational Inference (NRI) from dynamics prediction satisfies this condition, yet on a source-localisation task across 180 CFD-simulated hydrogen-leak scenarios and 180 acoustic scenarios, it degrades performance by 116% and 201% relative to a flexible, task-optimised attention baseline, while a physics-based prior (Green's function) degrades by only 69-72%. Four independent lines of evidence show this is not a tuning failure: NRI improves only 0.5% when given 18x more training data (versus 16.6% for the task-optimised baseline, p<0.001); performance is insensitive to the NRI edge threshold across a wide range; the dynamics-learned graph overlaps the task-optimal graph on only 6% of edges; and two further dynamics-derived structure estimators (correlation- and mutual-information-based) show no measurable benefit over a structure-free baseline, with the correlation-based estimator performing markedly worse. We formalise this gap as a statement about approximation error that the edge-accuracy condition cannot control, and we provide a lightweight transferability test (Jaccard similarity against a partially-observed target-task graph) that separates successful from failed transfer in all four domain/structure pairs we evaluate, using under an hour of computation and 15-20% of target-domain data; we present this as a heuristic calibrated on few cases, not a validated general threshold.
Figures & tables
Localiser
CFD RMSE (m) ↓
Acoustic RMSE (m) ↓
MLP (no structure)
14.57±0.24
4.35±0.07
Green’s Function
8.96±0.07
1.61±0.05
Learned Attention
5.21±0.22
0.95±0.04
NRI-Graph
11.26±0.10
2.86±0.05
Table 1: Structure comparison, n=15 sensors, mean ± std over 10 seeds. The ranking (Attention > Green’s > NRI > MLP) is preserved across two domains governed by different physics.
m
MLP
Green’s
Attention
NRI-Graph
10
15.94±1.28
8.63±0.38
6.26±0.54
11.32±0.65
180
14.57±0.24
8.96±0.07
5.21±0.22
11.26±0.10
Improvement
8.6%
−3.9%
16.6%
0.5%∗∗∗
Table 2: Data scaling: RMSE (m) as a function of training scenarios m (CFD domain). NRI-Graph’s improvement is an order of magnitude smaller than Attention’s despite identical data growth.
Domain
Structure
RMSE
Degradation
Est. J
CFD
Green’s Function
8.96 m
72%
≈0.31
NRI Learned Graph
11.26 m
116%
≈0.06
Acoustic
Green’s Function
1.61 m
69%
≈0.47
NRI Learned Graph
2.86 m
201%
≈0.19
Table 3: Transfer outcomes vs. estimated structural similarity to the task-optimal graph. A J≥0.30 threshold separates successful from failed transfer in all four cases.
Method
RMSE (m) ↓
Sparsity
Est. J
MLP (no structure)
14.57±0.24
—
0.00
NRI-Graph
11.26±0.10
97%
≈0.06
Correlation-Graph (best τ )
18.61±0.01
5%
≈1.00
Mutual-Information-Graph (best pct.)
14.48±0.00
80%
≈0.37
Table 4: Two additional dynamics-learned structure baselines (CFD domain, n=15 , mean ± std over 10 seeds), neither derived from NRI. Correlation-Graph clearly underperforms the unstructured MLP; Mutual-Information-Graph is statistically indistinguishable from it. Neither provides a measurable benefit, ruling out an NRI-architecture-specific explanation for why dynamics-derived structure fails to help.
CFD RMSE (m)
Acoustic RMSE (m)
n
MLP
NRI-Graph
MLP
NRI-Graph
10
14.02
12.37
3.60
3.43
20
15.17
10.53
5.19
2.41
Change
+8%
−15%
+44%
−30%
Table 5: Sensor-count scaling. The unstructured MLP degrades with more sensors in both domains; all structured methods improve, confirming that NRI-Graph’s hypothesis class is genuinely sparsity-constrained even though the constraint is misaligned with the task.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
CFD
Acoustic
Localiser
Train
Test
Gap
Train
Test
Gap
Learned Attention
5.45
5.67
+4.1%
1.05
1.15
+10.1%
Green’s Function
8.88
9.44
+6.3%
1.62
1.61
−0.7%
NRI-Graph
11.48
11.08
−3.4%
3.27
3.20
−2.0%
MLP (no structure)
15.39
16.92
+10.0%
4.78
5.03
+5.2%
Appendix
Table 6: Out-of-distribution generalisation. Green’s Function shows near-perfect acoustic robustness, consistent with its structure being derived directly from the governing physics rather than from data.
m
MLP
Green’s
Attention
NRI-Graph
10
15.94±1.28
8.63±0.38
6.26±0.54
11.32±0.65
25
15.54±0.46
8.02±0.28
5.26±0.39
11.32±0.29
50
17.84±0.47
9.61±0.21
5.76±0.21
12.10±0.27
100
16.68±0.38
9.35±0.17
5.75±0.16
11.58±0.19
150
15.08±0.20
9.25±0.15
5.45±0.17
11.61±0.15
180
14.57±0.24
8.96±0.07
5.21±0.22
11.26±0.10
Appendix
Table 7: Full data-scaling curve (CFD domain), RMSE (m) as a function of training scenarios m , mean ± std over 10 seeds.
Configuration
Avg. degree
Sparsity
RMSE (m)
τ=0.1
1.1
97%
11.50±0.07
τ∈[0.2,0.5]
1.0
97%
11.49±0.11
τ=0.7 (disconnected)
0.0
100%
17.42±0.13
Top-2 connectivity
2.0
94%
14.42±0.20
Top-4 connectivity
4.0
88%
13.79±0.16
Appendix
Table 8: NRI-Graph ablation over edge threshold τ and top- k connectivity (CFD domain). Performance is flat across τ∈[0.2,0.5] and degrades under denser top- k connectivity, ruling out a simple sparsity-tuning explanation.