A Response Theory Probe for Learned Stochastic AI Simulators, Tested on Lorenz-63
Authors: João Böger, Simon Driscoll, Niccolò Zagli, Valerio Lucarini, Francisco Camara Pereira
Organizations: Department of Technology, Management and Economics Technical University of Denmark Kongens Lyngby, 2800, DK · Department of Applied Mathematics and Theoretical Physics University of Cambridge Cambridge, CB3 0WA, UK · School of Computing and Mathematical Sciences University of Leicester Leicester, LE1 7RH, UK · School of Sciences Great Bay University Dongguan, P.R. China
Machine-learning emulators of chaotic and stochastic systems are usually validated on forecast skill and long-run statistics. Neither certifies that an emulator responds correctly to forcing, the property that projection and attribution studies rely on. Linear response theory makes this testable: the forced response follows from unperturbed correlations through a generalized fluctuation-dissipation relation, and decomposes over the stochastic Ruelle-Pollicott resonances of the Koopman generator. Building on the Koopmanism Response framework, we turn this into a calibrated, mode-resolved test for learned surrogates: each surrogate rollout passes or fails each check, and failure rates are compared with those of independent realizations of the true system. On stochastic Lorenz-63, a three-variable toy model, we evaluate SINDy, an MLP, a reservoir computer, a neural ODE and a neural SDE with learned diffusion, over up to 80 rollouts each. A sparse-regression model with the correct library passes every check at rates consistent with the true system. Invariant-statistics fidelity and response fidelity dissociate in both directions: a quarter of reservoir-computer rollouts pass every invariant-statistics check and match the static susceptibility χ(0), yet misrepresent the slow relaxation modes, while the neural ODE and SDE rarely meet the invariant-statistics floor but recover those modes in three quarters of rollouts. As expected of a time-integrated quantity dominated here by fast relaxation, χ(0) does not separate these cases. For a fixed network, the training formulation (one-step drift, flow map, or multi-step through the integrator) decides which of these properties it gets right.
Figures & tables
Figure 1: Invariant statistics, static susceptibility and resonances, per model. (a) x -marginal density: truth median and 5–95% range over 50 realizations; per model, the rollout with median χ(0) (MLP-drift not shown). (b) Accumulated response χ(t)=∫0tGx(s)ds , which tends to χ(0) : truth median and 5–95% range; per model, the median over rollouts; blue band: the χ(0) band of Table 1 . (c) Leading Koopman/RP resonances ( Imλ≥0 ): reference (stars) and, per model, the median of each leading group over rollouts, with its 5–95% range in Reλ ; real modes spread vertically for visibility. Dashed and dotted lines: the shortest-UPO twirling frequency ωp1 and the intra-lobe turnover frequency ωq1 [ 19 ] .
Invariant statistics
Slow ( Reλ )
Oscillatory ( ω )
Model
n
W1
ACF
switch
f+
χ(0)
−0.366
−0.508
8.42
9.85
9.98
16.68
threshold
≤ 0.284
≤ 0.089
0.985–1.035
0.487–0.517
0.162–0.309
≥ 0.945
≥ 0.922
≥ 0.990
≥ 0.968
≥ 0.960
≥ 0.990
Truth
50
0.076
0.038
1.011
0.500
0.224(0.034)
0.992
0.991
0.999
0.990
0.991
0.998
SINDy
10
0.092
0.038
1.013
0.494
0.215(0.029)
0.994
0.986
0.999
0.990
0.993
0.998
MLP-drift
80
5.599
3.724
0.008
0.477
1.045(1.278)
0.043
0.319
0.607
0.374
0.290
0.115
MLP-flow
80
2.193
0.735
0.725
0.474
0.636(0.313)
0.820
0.755
0.963
0.751
0.691
0.245
Table 1: Invariant statistics, χ(0) and eigenfunction overlaps, per model. Medians over scored rollouts; χ(0) as mean with the standard deviation in subscript. First row: the floor of each invariant statistic, the χ(0) band and the overlap threshold of each leading mode, labelled by its reference eigenvalue or frequency. Red: the median lies outside it. Bold: the best model in each column (lowest distance; switch ratio closest to one and f+ to one half; χ(0) mean closest to the truth’s; highest overlap). n : scored rollouts (RC: 2 diverged, left out).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Slow ( Reλ )
Oscillatory ( ω )
Model
n
-0.366
-0.508
8.42
9.85
9.98
16.68
merged
threshold
0.9450 (0.9450)
0.9216 (0.9216)
0.9900 (0.9984)
0.9681 (0.9681)
0.9601 (0.9601)
0.9900 (0.9960)
0.9900 (0.9974)
Truth
50
0.992 96%
0.991 94%
0.999 100%
0.990 98%
0.991 100%
0.998 100%
0.998 100%
SINDy
10
0.994 100%
0.986 100%
0.999 100%
0.990 90%
0.993 90%
0.998 100%
0.998 100%
MLP-drift
80
0.043 0%
0.319 0%
0.607 0%
0.373 0%
0.290 0%
0.115 0%
0.231 0%
MLP-flow
80
0.820 8%
0.754 5%
0.963 0%
0.751 0%
0.691 0%
0.245 0%
0.486 0%
Appendix
Table 2: Per-mode eigenfunction overlaps. Each cell: median overlap O (Eq. ( 13 )) over scored rollouts (top) and the percentage of rollouts that pass the mode (bottom); shading: redder = more rollouts fail. Oscillatory passes use the applied threshold clip(mean−3std,0.5,0.99) from the 16 calibration truths (first row; uncapped value in parentheses); slow modes use the uncapped threshold. † : more than half of the passes are decided by the 0.99 cap, i.e. would fail under the uncapped threshold. Merged: the joint subspace of the two near-degenerate pairs (9.85 and 9.98), with null 0.9982±0.0003 over the 16 calibration truths; reported here only. Bottom block: SINDy trained on 100 time units (App. C.1 ) and the true drift rolled out on the NODE/NSDE integrator paths (App. B.8 ). n : scored rollouts.
Figure 2: Per-mode passes and the dissociation map. (a) Fraction of rollouts passing each leading mode (slow modes uncapped, oscillatory modes matched directly); † : more than half of the passes are decided by the 0.99 cap. MLP-drift passes no mode in any rollout (not shown). (b) One point per scored rollout: e is the largest exceedance over the four invariant-statistics floors ( W1 and ACF as value over floor; switch ratio and f+ as distance from the floor interval’s centre over its half-width), m the smaller of the two slow-mode overlaps over their thresholds; e≤1 is inside every floor and m≥1 passes both slow modes. Hollow markers: χ(0) outside the band. The MLP models are not shown.
slow modes failed (%)
pass, rebuilt thresholds (%)
n
calibration set
rebuilt
λ≈−0.366
λ≈−0.508
Truth (perfect surrogate)
50
6
2
98
98
SINDy (full data)
10
0
0
100
100
MLP-drift
80
100
100
0
0
MLP-flow
80
97
94
10
19
RC
78
76
55
45
62
Appendix
Table 3: Slow-mode threshold sensitivity. The two slow-mode thresholds rebuilt from the 50 scored fresh truths by a leave-one-out minimum rule: each truth is judged against the minimum overlap of the other truths (its row is the leave-one-out false-alarm rate), every model against the minimum over all truths ( λ≈−0.366 : 0.9257 instead of 0.9450; λ≈−0.508 : 0.8693 instead of 0.9216). The stored per-mode overlaps are re-thresholded; nothing is re-run, and Table 8 keeps the calibration-set thresholds. Rollouts inside all four invariant-statistics floors and the χ(0) band that fail the slow modes under the rebuilt thresholds: RC 14 of 26, truth 1 of 48 (Fisher exact p=2.6×10−7 ).
n
leading share
leading ∣ share ∣
accumulated by t = 1
reference
1
0.34
–
0.65
truth
10
0.33 [-0.15, 0.60]
0.06
0.89 [0.60, 1.32]
SINDy-100
10
-0.06 [-0.52, 0.40]
0.04
0.97 [0.76, 1.30]
RC
10
0.29 [-0.06, 0.60]
0.07
0.87 [0.65, 1.65]
NODE
10
0.24 [-0.36, 0.50]
0.04
0.87 [0.62, 1.54]
NSDE
10
0.08 [-0.41, 0.57]
0.03
0.87 [0.62, 1.51]
Appendix
Table 4: Where χ(0) sits. Signed share of χ(0) carried by the leading resonances (median [min, max]), their share of the summed absolute mode contributions, and the fraction of χ(0) accumulated by t=1 .
W1
ACF
switch
f+
all four
band (E1b)
band (pooled)
leave-one-out out-rate
0.02
0.02
0.04
0.04
0.10
0.04
0.02
runner truths outside
0
0
2
0
–
6
–
one-sided binomial p
1
1
0.6
1
–
0.0144
–
Appendix
Table 5: Null calibration. Leave-one-out out-rates on the independent true ensemble (floors, E1b band) and the pooled band; out-of-range counts of the 50 runner truths and the one-sided binomial test against the leave-one-out rate. The E1b band test fails ( p<0.05 ), so the band is the pooled range [0.162,0.309] over 100 truths; no floor test fails, so all four floors stay as built. Runner and E1b χ(0) variances: F=1.46 , p=0.191 .
Model
n
pooled [0.162,0.309]
E1b [0.163,0.266]
N = 18 [0.18,0.324]
runner [0.162,0.309]
diverged
Truth (perfect surrogate)
50
50
44
44
50
0
SINDy (full data)
10
10
10
9
10
0
MLP-drift
80
7
5
8
7
0
MLP-flow
80
1
0
2
1
0
RC
78
72
67
59
72
2
NODE
80
73
61
66
73
0
Appendix
Table 6: χ(0) in-band counts under each band (sensitivity rows), and diverged rollouts (non-finite or a coordinate beyond 1.5× its true maximum: ∣x∣,∣y∣,∣z∣>31.7,43.6,76.1 ), excluded from all rates.
n
invariant (%)
χ(0) (%)
slow (%)
oscill. (%)
χ(0) median
shift
Truth
50
4
0
6
2
0.221
–
G0, NODE path
10
100
0
10
40
0.265
0.0443
G0, NSDE path
10
100
30
0
40
0.275
0.0549
Appendix
Table 7: Discretization sensitivity (G0) : the true drift through the NODE/NSDE RK4 step h , failure rates per rung and the shift of the median χ(0) from the truth median.
Invariant statistics (%)
χ(0)
Modes (%)
Model
n (div.)
any
W1
ACF
switch
f+
mean (sd)
out (%)
slow
oscill.
Truth
50 (0)
4
0
0
4
0
0.224(0.034)
0 (2)
6
2
SINDy
10 (0)
20
10
0
0
20
0.215(0.029)
0
0
10
MLP-drift
80 (0)
100
100
100
100
96
1.045(1.278)
91
100
100
MLP-flow
80 (0)
100
100
93
100
100
0.636(0.313)
99
98
100
RC
78 (2)
63
3
27
58
6
0.203(0.032)
8
76
100
Appendix
Table 8: Per-rollout failure rates , in percent of scored rollouts (redder = fails more often); a rollout fails a rung (invariant statistics, χ(0) , slow modes, oscillatory modes) if it fails any check in it. Truth: fresh true realizations scored as a perfect surrogate, whose failure rate is the false-alarm rate; SINDy: trained on the full record. Invariant statistics: any of the four checks, then each check alone, outside its floor (the maximum over the 50 independent true realizations scored against the reference seed 101): W1≤0.284 , ACF distance ≤0.0885 , switch ratio in [0.985,1.03] , f+ in [0.487,0.517] . χ(0) : mean with the standard deviation in subscript; fails outside [0.162,0.309] , the range of the 100 pooled true realizations (50 independent-IC plus 50 scored through the runner); the truth row is inside by construction, and its leave-one-out out-rate is given in parentheses. Slow modes: the two real reference modes ( λ≈−0.366 and −0.508 ), overlap below the uncapped threshold mean −3 std over 16 calibration truths ( 0.9450 , 0.9216 ). Oscillatory modes: the pairs at ω≈8.42 , 9.85 , 9.98 , 16.68 , matched directly with thresholds clip(mean−3std,0.5,0.99) ; the 8.42 and 16.68 thresholds sit at the cap. MLP-drift: 33 of 80 rollouts collapse onto one wing (scored). n counts scored rollouts; diverged rollouts (non-finite, or a coordinate beyond 1.5× its maximum over the true realizations) are left out of the rates: RC 2. Transformer: exploratory, fixed configuration, not tuned (App. C.5 ).
Table 9: Parameters. System, instrument, seeds (top) and per-model training and rollout (bottom). Every non-NSDE surrogate is rolled out with the true noise amplitude s (oracle noise).
term
true
SINDy-100
SINDy (full)
σ ( y in x˙ )
10.000
10.520
10.001
− ( x in x˙ )
10.000
10.755
10.003
ρ ( x in y˙ )
28.000
27.734
27.967
β ( −z in z˙ )
2.667
2.402
2.667
const in x˙
0.000
-3.423
0.000
const in y˙
0.000
3.139
0.307
Appendix
Table 10: SINDy coefficients : true, SINDy-100 and SINDy on the full training span (the same recipe; large terms are those above the STLSQ threshold).
x
y
z
SINDy (full data)
0.014
0.070
0.043
SINDy-100
0.717
0.537
0.647
NODE (median over 8 fits)
0.577
1.125
0.961
MLP-drift
2.048
4.399
4.213
secant bias of the MLP-drift target (exact)
2.192
4.264
4.541
secant bias, leading order (Δt/2)JFF
2.439
4.740
5.056
Appendix
Table 11: Drift RMSE against the true drift on the 20000-point grid drawn from the held-out true realization (seed 909).
L
lr
MSE 1
MSE 5
MSE 20
epochs
train (s)
drift RMSE x/y/z
5
0.001
0.547
1.543
6.225
25
177
0.490/0.921/0.877
5
0.0003
0.548
1.547
6.242
25
177
0.648/1.078/0.986
20
0.001
0.549
1.555
6.268
14
342
0.841/1.483/1.479
20
0.0003
0.548
1.551
6.239
25
601
0.800/1.229/1.260
Appendix
Table 12: NODE sweep (one seed): validation MSE ( ×10−3 , standardized) at horizons L=1,5,20 ; selection by the lowest MSE at the common horizon L=20 (first row). Drift RMSE on the grid is reported only.
x
y
z
NSDE mean σϕ/s (median of 80)
0.9587
0.9990
0.9883
NSDE constant σ^/s
0.9567
0.9990
0.9891
true drift, step h , constant σ^/s
0.9567
0.9982
0.9885
(1−e−2ah)/(2ah)
0.9520
0.9950
0.9868
Appendix
Table 13: NSDE learned noise relative to the true s , at h=0.01 ; a=(σ,1,β) .
Apr 22, 2026·Gabriel Melo, Leonardo Santiago, Peter Y. LuEmulatorsChaos
LTCI, Télécom Paris, Institut Polytechnique de Paris, Palaiseau · Department of Mechanical and Aerospace Engineering, North Carolina State University, Raleigh, NC · University of Campinas +1