A Response Theory Probe for Learned Stochastic AI Simulators, Tested on Lorenz-63
Authors: João Böger, Simon Driscoll, Niccolò Zagli, Valerio Lucarini, Francisco Camara Pereira
Organizations: Department of Technology, Management and Economics Technical University of Denmark Kongens Lyngby, 2800, DK · Department of Applied Mathematics and Theoretical Physics University of Cambridge Cambridge, CB3 0WA, UK · School of Computing and Mathematical Sciences University of Leicester Leicester, LE1 7RH, UK · School of Sciences Great Bay University Dongguan, P.R. China
Machine-learning emulators of chaotic and stochastic systems are usually validated on forecast skill and long-run statistics. Neither certifies that an emulator responds correctly to forcing, the property that projection and attribution studies rely on. Linear response theory makes this testable: the forced response follows from unperturbed correlations through a generalized fluctuation-dissipation relation, and decomposes over the stochastic Ruelle-Pollicott resonances of the Koopman generator. Building on the Koopmanism Response framework, we turn this into a calibrated, mode-resolved test for learned surrogates: each surrogate rollout passes or fails each check, and failure rates are compared with those of independent realizations of the true system. On stochastic Lorenz-63, a three-variable toy model, we evaluate SINDy, an MLP, a reservoir computer, a neural ODE and a neural SDE with learned diffusion, over up to 80 rollouts each. A sparse-regression model with the correct library passes every check at rates consistent with the true system. Invariant-statistics fidelity and response fidelity dissociate in both directions: a quarter of reservoir-computer rollouts pass every invariant-statistics check and match the static susceptibility χ(0), yet misrepresent the slow relaxation modes, while the neural ODE and SDE rarely meet the invariant-statistics floor but recover those modes in three quarters of rollouts. As expected of a time-integrated quantity dominated here by fast relaxation, χ(0) does not separate these cases. For a fixed network, the training formulation (one-step drift, flow map, or multi-step through the integrator) decides which of these properties it gets right.
Figures & tables
Figure 1: Invariant statistics, static susceptibility and resonances, per model. (a) x -marginal density: truth median and 5–95% range over 50 realizations; per model, the rollout with median χ(0) (MLP-drift not shown). (b) Accumulated response χ(t)=∫0tGx(s)ds , which tends to χ(0) : truth median and 5–95% range; per model, the median over rollouts; blue band: the χ(0) band of Table 1 . (c) Leading Koopman/RP resonances ( Imλ≥0 ): reference (stars) and, per model, the median of each leading group over rollouts, with its 5–95% range in Reλ ; real modes spread vertically for visibility. Dashed and dotted lines: the shortest-UPO twirling frequency ωp1 and the intra-lobe turnover frequency ωq1 [ 19 ] .
Invariant statistics
Slow ( Reλ )
Oscillatory ( ω )
Model
n
W1
ACF
switch
f+
χ(0)
−0.366
−0.508
8.42
9.85
9.98
16.68
threshold
≤ 0.284
≤ 0.089
0.985–1.035
0.487–0.517
0.162–0.309
≥ 0.945
≥ 0.922
≥ 0.990
≥ 0.968
≥ 0.960
≥ 0.990
Truth
50
0.076
0.038
1.011
0.500
0.224(0.034)
0.992
0.991
0.999
0.990
0.991
0.998
SINDy
10
0.092
0.038
1.013
0.494
0.215(0.029)
0.994
0.986
0.999
0.990
0.993
0.998
MLP-drift
80
5.599
3.724
0.008
0.477
1.045(1.278)
0.043
0.319
0.607
0.374
0.290
0.115
MLP-flow
80
2.193
0.735
0.725
0.474
0.636(0.313)
0.820
0.755
0.963
0.751
0.691
0.245
Table 1: Invariant statistics, χ(0) and eigenfunction overlaps, per model. Medians over scored rollouts; χ(0) as mean with the standard deviation in subscript. First row: the floor of each invariant statistic, the χ(0) band and the overlap threshold of each leading mode, labelled by its reference eigenvalue or frequency. Red: the median lies outside it. Bold: the best model in each column (lowest distance; switch ratio closest to one and f+ to one half; χ(0) mean closest to the truth’s; highest overlap). n : scored rollouts (RC: 2 diverged, left out).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Slow ( Reλ )
Oscillatory ( ω )
Model
n
-0.366
-0.508
8.42
9.85
9.98
16.68
merged
threshold
0.9450 (0.9450)
0.9216 (0.9216)
0.9900 (0.9984)
0.9681 (0.9681)
0.9601 (0.9601)
0.9900 (0.9960)
0.9900 (0.9974)
Truth
50
0.992 96%
0.991 94%
0.999 100%
0.990 98%
0.991 100%
0.998 100%
0.998 100%
SINDy
10
0.994 100%
0.986 100%
0.999 100%
0.990 90%
0.993 90%
0.998 100%
0.998 100%
MLP-drift
80
0.043 0%
0.319 0%
0.607 0%
0.373 0%
0.290 0%
0.115 0%
0.231 0%
MLP-flow
80
0.820 8%
0.754 5%
0.963 0%
0.751 0%
0.691 0%
0.245 0%
0.486 0%
Appendix
Table 2: Per-mode eigenfunction overlaps. Each cell: median overlap O (Eq. ( 13 )) over scored rollouts (top) and the percentage of rollouts that pass the mode (bottom); shading: redder = more rollouts fail. Oscillatory passes use the applied threshold clip(mean−3std,0.5,0.99) from the 16 calibration truths (first row; uncapped value in parentheses); slow modes use the uncapped threshold. † : more than half of the passes are decided by the 0.99 cap, i.e. would fail under the uncapped threshold. Merged: the joint subspace of the two near-degenerate pairs (9.85 and 9.98), with null 0.9982±0.0003 over the 16 calibration truths; reported here only. Bottom block: SINDy trained on 100 time units (App. C.1 ) and the true drift rolled out on the NODE/NSDE integrator paths (App. B.8 ). n : scored rollouts.
Figure 2: Per-mode passes and the dissociation map. (a) Fraction of rollouts passing each leading mode (slow modes uncapped, oscillatory modes matched directly); † : more than half of the passes are decided by the 0.99 cap. MLP-drift passes no mode in any rollout (not shown). (b) One point per scored rollout: e is the largest exceedance over the four invariant-statistics floors ( W1 and ACF as value over floor; switch ratio and f+ as distance from the floor interval’s centre over its half-width), m the smaller of the two slow-mode overlaps over their thresholds; e≤1 is inside every floor and m≥1 passes both slow modes. Hollow markers: χ(0) outside the band. The MLP models are not shown.
slow modes failed (%)
pass, rebuilt thresholds (%)
n
calibration set
rebuilt
λ≈−0.366
λ≈−0.508
Truth (perfect surrogate)
50
6
2
98
98
SINDy (full data)
10
0
0
100
100
MLP-drift
80
100
100
0
0
MLP-flow
80
97
94
10
19
RC
78
76
55
45
62
Appendix
Table 3: Slow-mode threshold sensitivity. The two slow-mode thresholds rebuilt from the 50 scored fresh truths by a leave-one-out minimum rule: each truth is judged against the minimum overlap of the other truths (its row is the leave-one-out false-alarm rate), every model against the minimum over all truths ( λ≈−0.366 : 0.9257 instead of 0.9450; λ≈−0.508 : 0.8693 instead of 0.9216). The stored per-mode overlaps are re-thresholded; nothing is re-run, and Table 8 keeps the calibration-set thresholds. Rollouts inside all four invariant-statistics floors and the χ(0) band that fail the slow modes under the rebuilt thresholds: RC 14 of 26, truth 1 of 48 (Fisher exact p=2.6×10−7 ).
n
leading share
leading ∣ share ∣
accumulated by t = 1
reference
1
0.34
–
0.65
truth
10
0.33 [-0.15, 0.60]
0.06
0.89 [0.60, 1.32]
SINDy-100
10
-0.06 [-0.52, 0.40]
0.04
0.97 [0.76, 1.30]
RC
10
0.29 [-0.06, 0.60]
0.07
0.87 [0.65, 1.65]
NODE
10
0.24 [-0.36, 0.50]
0.04
0.87 [0.62, 1.54]
NSDE
10
0.08 [-0.41, 0.57]
0.03
0.87 [0.62, 1.51]
Appendix
Table 4: Where χ(0) sits. Signed share of χ(0) carried by the leading resonances (median [min, max]), their share of the summed absolute mode contributions, and the fraction of χ(0) accumulated by t=1 .
W1
ACF
switch
f+
all four
band (E1b)
band (pooled)
leave-one-out out-rate
0.02
0.02
0.04
0.04
0.10
0.04
0.02
runner truths outside
0
0
2
0
–
6
–
one-sided binomial p
1
1
0.6
1
–
0.0144
–
Appendix
Table 5: Null calibration. Leave-one-out out-rates on the independent true ensemble (floors, E1b band) and the pooled band; out-of-range counts of the 50 runner truths and the one-sided binomial test against the leave-one-out rate. The E1b band test fails ( p<0.05 ), so the band is the pooled range [0.162,0.309] over 100 truths; no floor test fails, so all four floors stay as built. Runner and E1b χ(0) variances: F=1.46 , p=0.191 .
Model
n
pooled [0.162,0.309]
E1b [0.163,0.266]
N = 18 [0.18,0.324]
runner [0.162,0.309]
diverged
Truth (perfect surrogate)
50
50
44
44
50
0
SINDy (full data)
10
10
10
9
10
0
MLP-drift
80
7
5
8
7
0
MLP-flow
80
1
0
2
1
0
RC
78
72
67
59
72
2
NODE
80
73
61
66
73
0
Appendix
Table 6: χ(0) in-band counts under each band (sensitivity rows), and diverged rollouts (non-finite or a coordinate beyond 1.5× its true maximum: ∣x∣,∣y∣,∣z∣>31.7,43.6,76.1 ), excluded from all rates.
n
invariant (%)
χ(0) (%)
slow (%)
oscill. (%)
χ(0) median
shift
Truth
50
4
0
6
2
0.221
–
G0, NODE path
10
100
0
10
40
0.265
0.0443
G0, NSDE path
10
100
30
0
40
0.275
0.0549
Appendix
Table 7: Discretization sensitivity (G0) : the true drift through the NODE/NSDE RK4 step h , failure rates per rung and the shift of the median χ(0) from the truth median.
Invariant statistics (%)
χ(0)
Modes (%)
Model
n (div.)
any
W1
ACF
switch
f+
mean (sd)
out (%)
slow
oscill.
Truth
50 (0)
4
0
0
4
0
0.224(0.034)
0 (2)
6
2
SINDy
10 (0)
20
10
0
0
20
0.215(0.029)
0
0
10
MLP-drift
80 (0)
100
100
100
100
96
1.045(1.278)
91
100
100
MLP-flow
80 (0)
100
100
93
100
100
0.636(0.313)
99
98
100
RC
78 (2)
63
3
27
58
6
0.203(0.032)
8
76
100
Appendix
Table 8: Per-rollout failure rates , in percent of scored rollouts (redder = fails more often); a rollout fails a rung (invariant statistics, χ(0) , slow modes, oscillatory modes) if it fails any check in it. Truth: fresh true realizations scored as a perfect surrogate, whose failure rate is the false-alarm rate; SINDy: trained on the full record. Invariant statistics: any of the four checks, then each check alone, outside its floor (the maximum over the 50 independent true realizations scored against the reference seed 101): W1≤0.284 , ACF distance ≤0.0885 , switch ratio in [0.985,1.03] , f+ in [0.487,0.517] . χ(0) : mean with the standard deviation in subscript; fails outside [0.162,0.309] , the range of the 100 pooled true realizations (50 independent-IC plus 50 scored through the runner); the truth row is inside by construction, and its leave-one-out out-rate is given in parentheses. Slow modes: the two real reference modes ( λ≈−0.366 and −0.508 ), overlap below the uncapped threshold mean −3 std over 16 calibration truths ( 0.9450 , 0.9216 ). Oscillatory modes: the pairs at ω≈8.42 , 9.85 , 9.98 , 16.68 , matched directly with thresholds clip(mean−3std,0.5,0.99) ; the 8.42 and 16.68 thresholds sit at the cap. MLP-drift: 33 of 80 rollouts collapse onto one wing (scored). n counts scored rollouts; diverged rollouts (non-finite, or a coordinate beyond 1.5× its maximum over the true realizations) are left out of the rates: RC 2. Transformer: exploratory, fixed configuration, not tuned (App. C.5 ).
Table 9: Parameters. System, instrument, seeds (top) and per-model training and rollout (bottom). Every non-NSDE surrogate is rolled out with the true noise amplitude s (oracle noise).
term
true
SINDy-100
SINDy (full)
σ ( y in x˙ )
10.000
10.520
10.001
− ( x in x˙ )
10.000
10.755
10.003
ρ ( x in y˙ )
28.000
27.734
27.967
β ( −z in z˙ )
2.667
2.402
2.667
const in x˙
0.000
-3.423
0.000
const in y˙
0.000
3.139
0.307
Appendix
Table 10: SINDy coefficients : true, SINDy-100 and SINDy on the full training span (the same recipe; large terms are those above the STLSQ threshold).
x
y
z
SINDy (full data)
0.014
0.070
0.043
SINDy-100
0.717
0.537
0.647
NODE (median over 8 fits)
0.577
1.125
0.961
MLP-drift
2.048
4.399
4.213
secant bias of the MLP-drift target (exact)
2.192
4.264
4.541
secant bias, leading order (Δt/2)JFF
2.439
4.740
5.056
Appendix
Table 11: Drift RMSE against the true drift on the 20000-point grid drawn from the held-out true realization (seed 909).
L
lr
MSE 1
MSE 5
MSE 20
epochs
train (s)
drift RMSE x/y/z
5
0.001
0.547
1.543
6.225
25
177
0.490/0.921/0.877
5
0.0003
0.548
1.547
6.242
25
177
0.648/1.078/0.986
20
0.001
0.549
1.555
6.268
14
342
0.841/1.483/1.479
20
0.0003
0.548
1.551
6.239
25
601
0.800/1.229/1.260
Appendix
Table 12: NODE sweep (one seed): validation MSE ( ×10−3 , standardized) at horizons L=1,5,20 ; selection by the lowest MSE at the common horizon L=20 (first row). Drift RMSE on the grid is reported only.
x
y
z
NSDE mean σϕ/s (median of 80)
0.9587
0.9990
0.9883
NSDE constant σ^/s
0.9567
0.9990
0.9891
true drift, step h , constant σ^/s
0.9567
0.9982
0.9885
(1−e−2ah)/(2ah)
0.9520
0.9950
0.9868
Appendix
Table 13: NSDE learned noise relative to the true s , at h=0.01 ; a=(σ,1,β) .
Many scientific systems exhibit uncertainty from stochastic forcing, unresolved degrees of freedom, or imperfect observations, making reliable surrogate forecasting fundamentally distributional rather than pointwise. For such systems, deterministic neural surrogates fail to capture statistical measures and forecast uncertainty. We introduce TRIE, an evaluation framework for stochastic PDE surrogates that asks whether models reproduce invariant measures, provide trustworthy predictive uncertainty, and scale to efficient probabilistic generation. We demonstrate TRIE on two stationary chaotic spatially extended SPDEs, stochastic Kuramoto--Sivashinsky and stochastic Kolmogorov flow, across 11 parameter values. Our evaluation shows that standard pointwise-trained neural surrogates can produce plausible short rollouts while failing to match long-time statistical structure. Approximate uncertainty methods such as Monte Carlo dropout and heteroscedastic Gaussian likelihoods produce stochastic forecasts, but are often miscalibrated and overconfident under temporal and spatial uncertainty diagnostics. Across these criteria, generative models provide the most consistent performance, accurately capturing invariant measure statistics and achieving the lowest CRPS in all reported probabilistic settings. Finally, we show that latent generative models with automatic dimension discovery retain much of this statistical fidelity while reducing Kolmogorov inference time by roughly 12×. We release our code and data at https://github.com/scailab/TRIE-SPDE-Bench to support reproducible evaluation of stochastic PDE forecasting models.
Bharat Srikishan, Javier E. Santos, Nikhil Muralidhar +1
Stevens Institute of Technology · Los Alamos National Laboratory
Chaos arises in many complex dynamical systems, from weather to power grids, but is difficult to accurately model with data-driven methods such as machine learning emulators. While emulators are promising tools for accelerating simulations and solving inverse problems, they still struggle to learn chaotic dynamics, where sensitivity to initial conditions renders exact long-term forecasts infeasible, especially given noisy data. Recent work instead trains emulators to match the statistical properties of chaotic attractors, but these approaches often rely on handcrafted summary statistics or large, diverse multi-environment datasets. In this work, we propose a family of adversarial optimal transport objectives that can jointly learn high-quality summary statistics and a physically consistent emulator from a single noisy trajectory. We theoretically analyze and experimentally validate a Sinkhorn divergence formulation (2-Wasserstein) and a WGAN-style dual formulation (1-Wasserstein) of our approach. Numerical experiments across a variety of chaotic systems, including ones with high-dimensional spatiotemporal chaos, show that emulators trained using our proposed objectives have significantly improved long-term statistical fidelity.
Gabriel Melo, Leonardo Santiago, Peter Y. Lu
LTCI, Télécom Paris, Institut Polytechnique de Paris, Palaiseau · Department of Mechanical and Aerospace Engineering, North Carolina State University, Raleigh, NC · University of Campinas +1
Large language models (LLMs) can fluently generate student-like responses, making them attractive as simulated students for training and evaluating AI tutors and human educators. Yet such simulators are typically evaluated by output similarity to real students, not by whether they behave like students with coherent misconceptions during interaction. We introduce a controlled framework for evaluating misconception faithfulness, whether a simulator maintains a misconception-driven belief state and updates selectively when feedback addresses the underlying misconception. Central to our framework is a misconception-contrastive feedback protocol that compares targeted feedback against two controls: misaligned feedback (targeting a different but plausible misconception) and generic feedback (only identifying answer is wrong). We propose Selective Flip Score (SFS), which quantifies how much more often a simulator flips its answer under targeted feedback than under contrastive controls. Across seven LLMs (4B-120B), multiple datasets, and prompting strategies, simulators exhibit near-zero SFS, correcting their answers at similarly high rates regardless of feedback relevance. Further analyses reveal a sycophantic failure mode: models behave less like students with misconceptions but more like problem-solvers who treat any corrective signal as a cue to abandon the simulated belief and re-solve from internal knowledge. To address this, we develop a post-training pipeline spanning supervised fine-tuning (SFT), preference optimization, and reinforcement learning (RL) with an SFS-aligned reward; SFT yields notable gains up to +0.56, and SFS-aligned RL provides more consistent improvements than preference optimization. Our results establish misconception faithfulness as a challenging yet trainable property, motivating a shift from static output matching toward interactive, belief-aware student modeling.
Heejin Do, Shashank Sonkar, Mrinmaya Sachan
ETH Zürich, ETH AI Center · University of Central Florida · ETH Zürich