Scientific dynamics forecasting is often framed as an architecture choice, although deployment is also determined by observed history, rollout feedback, compute budget, physical objective, and test distribution. We formulate protocol-dependent model selection and introduce ProtocolMatch, a compute-matched, validation-selected, and failure-preserving evaluation framework. On driven quantum-spin dynamics, we compare recurrent, patched-attention, causal-attention, and low-rank linear predictors across three independently generated datasets. The causal-attention--recurrence ordering reverses as the training set grows within a fixed two-spin task, while a linear predictor has the lowest mean error in the six-spin local-observable comparison. Restricting observed history worsens every refreshed-history view but improves every closed-loop view in the four-spin study. A latest-state MLP has lower error than persistence on every dataset under state refresh across all five cells, yet its closed-loop rank varies by system and includes finite explosive errors. Physical penalties improve targeted consistency without reliably improving prediction error, and in-distribution intervals lose most coverage after a driving-frequency shift. Thus scientific model selection should return a predictor with its protocol and report accuracy, physical validity, and shifted-distribution reliability separately.
Figures & tables
Figure 1: The forecasting protocol determines both the information available to a predictor and how that predictor is compared. (a) Observed-history and closed-loop evaluation keep the predictor, future controls, and target fixed; only the source of the next state-history window changes. (b) ProtocolMatch pairs that declared task with controlled candidates, validation-only selection, independent-dataset aggregation, and failure-preserving reliability reports.
Factor
Values
Question
History
1, 5, 20
How much past state is useful?
Channels
Full, local
Which observables are available?
Rollout
Refresh, closed loop
Is feedback autonomous?
Clipping
Off, on
Are bounds imposed at inference?
Budget
30 s, 120 s
Is optimization time matched?
Test family
ID, freq., state shifts
Where must the model generalize?
Table 1: Protocol factors varied in the evaluation grid.
Figure 2: Replicated ranking reversals under matched optimization time. (a) LSTM-minus-causal error for closed-loop extrapolation changes sign between 64 and 256 training trajectories. (b) Local-minus-full-history error changes sign between observed-history and closed-loop rollout while the full prediction target remains fixed. Colored lines are independent-dataset means over three training seeds; the dashed black line is their average. Both panels use 20-step histories, a 120-second budget, identity-distribution tests, and no output clipping.
System and target
Train traj.
LSTM
PatchTST
Causal Transformer
Low-rank linear
Two-spin Ising, full Pauli
64
0.09402
0.65869
0.08681
0.19583
Two-spin Ising, full Pauli
256
0.02456
0.66745
0.04654
0.18849
Four-spin Ising, full Pauli
256
0.10608
0.12416
0.10902
0.12220
Two-spin XXZ, full Pauli
256
0.03074
0.25047
0.04540
0.18421
Six-spin Ising, local Pauli
256
0.08833
0.08835
0.07878
0.07668
Table 2: Nonidentity-observable MSE under unclipped closed-loop extrapolation. Bold marks the lowest mean within each row. Two/four-spin targets are full Pauli vectors; the six-spin target contains local observables only.
Observed-history refresh
Unclipped closed loop
System and target
Train traj.
Current-state MLP
Persistence
Current-state MLP
Persistence / direction
Two-spin Ising, full Pauli
64
0.00524±0.00094
0.09979±0.00323
0.49307±0.31782
0.34285±0.00661 / mixed
Two-spin Ising, full Pauli
256
0.00200±0.00038
0.09979±0.00323
0.22436±0.03245
0.34285±0.00661 / MLP 3/3
Four-spin Ising, full Pauli
256
0.03675±0.00150
0.04238±0.00158
2.73×1035±4.72×1035
0.11203±0.00119 / mixed
Two-spin XXZ, full Pauli
256
0.00630±0.00306
0.12930±0.00878
4.12985±5.75518
0.33538±0.00844 / persistence 3/3
Six-spin Ising, local Pauli
256
0.01362±0.00148
0.03222±0.00299
0.08737±0.00167
0.29335±0.00579 / MLP 3/3
Table 3: Latest-state controls on identity-distribution extrapolation. Entries are nonidentity MSE mean ± sample SD across three independent dataset means; MLP training seeds are averaged within each dataset first. The MLP uses its validation-selected 120-second checkpoint. “3/3” denotes the same paired direction in all datasets.
Figure 3: Accuracy, physical validity, and uncertainty are distinct outcomes. (a) A trace penalty usually lowers trace error but often raises prediction MSE for a two-spin LSTM; the 32 correlated views share fitted models. (b) Intervals calibrated on 64 independent identity-distribution trajectories approach nominal 90% coverage in distribution but lose most coverage after a frequency shift. Widths are full symmetric interval widths.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Evaluations
Independent datasets
Primary purpose
Core matched comparisons
288
3 per system cell
Architecture and data-regime ranks
History and input controls
540
3 per condition
Information–rollout interaction
Constraint terms and weights
216
1
Accuracy–validity trade-off
Six-spin extension
72
3
Local-observable scaling check
Primary matrix total
1,116
—
—
Diagnostic baselines (separate)
105
3 per system cell
Latest-state and persistence controls
Appendix
Table 4: Experimental inventory. Counts denote fitted architecture/configuration instances except for the diagnostic row, which includes 90 MLP fits and 15 deterministic persistence evaluations.
Contrast
Condition
Dataset 1
Dataset 2
Dataset 3
Mean
LSTM minus causal
64 trajectories
+0.01292
+0.00508
+0.00363
+0.00721
LSTM minus causal
256 trajectories
-0.02696
-0.01586
-0.02314
-0.02198
Local minus full input
Observed-history
+0.01967
+0.02011
+0.02162
+0.02046
Local minus full input
Closed-loop
-0.03221
-0.03072
-0.02776
-0.03023
Appendix
Table 5: Dataset-level contrasts in Figure 2 . Each dataset column averages three training seeds; the final column averages datasets. Positive architecture differences favor causal attention, and positive input differences favor full history.
System
Comparison
Lower-error views
All-data agreement
Two-spin Ising
LSTM history 5 vs. 20
27/32
9/32
Four-spin Ising
LSTM history 5 vs. 20
30/32
28/32
Four-spin Ising
Local vs. full, observed history
0/16
16/16 worse
Four-spin Ising
Local vs. full, closed loop
16/16
16/16 better
Appendix
Table 6: Direction summaries across the 32 correlated evaluation views. “All-data agreement” counts views in which the sign is the same for all three independently generated datasets.
Penalty weight
Views with lower trace error
Views with higher prediction MSE
0.001
31/32
26/32
0.01
32/32
17/32
0.1
32/32
21/32
Appendix
Table 7: Two-spin LSTM trace-penalty direction counts on the prespecified constraint dataset.
Predictor
ID coverage
Shift coverage
Width
LSTM
87.50%
28.13%
1.2767
PatchTST
91.32%
26.91%
7.6545
Causal Transformer
86.98%
16.49%
1.7195
Low-rank linear
87.67%
24.31%
2.3278
Appendix
Table 8: Nominal 90% whole-trajectory-region interval results for the two-spin Ising, 256-trajectory, full-history condition. Width is the full symmetric interval width.
Figure 4: Closed-loop amplification for one prespecified PatchTST model–trajectory pair. Each point is the maximum absolute model output in a successive 20-step block; the vertical axis is logarithmic. Observed-history refresh stays bounded, whereas unclipped feedback reaches 5.58×1023 before the next block becomes nonfinite (cross). This numerical diagnostic is not a population failure-rate estimate.
Scientific forecasting typically relies on direct state prediction, an approach that grows brittle under data scarcity, extended horizons, non-stationary dynamics, or high-dimensional complexity. While raw state trajectories are highly sensitive in these regimes, underlying local evolution rules often exhibit robust reusability. We introduce mechanism learning, a framework that forecasts future states by estimating the currently active local mechanism. Our method compresses local spatiotemporal fragments into mechanism descriptors, forming a data-driven, structured mechanism space where proximity reflects similar local evolution rules. To ground these estimates in observed data, we utilize prototype anchors, a set of representative mechanisms that sparsely cover the space of local rules. We evaluate this approach on Burgers dynamics, WeatherBench2, and Lorenz96. Empirically, the learned mechanism spaces resist collapse and maintain strong local consistency. Compared to direct prediction and other models including FNO, NODE, LSTM, and reservoir-family methods, our framework demonstrates predictive gains in fragile regimes: it significantly improves switching stability in Burgers dynamics and achieves state-of-the-art performance both under the scarce-data fixed-horizon WeatherBench2 protocol and in intermediate-complexity Lorenz96. Ablation studies and drift diagnostics confirm that these improvements are driven by finite prototype anchoring rather than sheer latent capacity. Together, these results establish mechanism learning as a principled, robust alternative to direct state prediction in forecasting complex systems.
Qian Jiang, Liping Sun
School of Computing · The Australian National University · Canberra, ACT 2601, Australia +3
Machine-learning models for physical systems are currently evaluated primarily through errors between predicted and reference states and, increasingly, through tests of physical consistency. These metrics assess whether predictions are accurate and satisfy selected physical requirements, but provide limited insight into whether learned trajectories reproduce the underlying dynamics. Domain experts examine such relationships through physical representations that expose relevant processes, interactions, and responses, but these analyses are often separated from typical machine-learning evaluation. We introduce a practical framework for evaluating dynamical fidelity through predictive structure in physical representation spaces. Experts define the representations, while reference trajectories determine which relationships are predictive and retained as evaluation tests. We demonstrate the approach in atmospheric forecasting using ERA5 representations of planetary-wave activity and Northern Annular Mode evolution, and evaluate Pangu-Weather, GraphCast, and FengWu. The models exhibit distinct departures from reference predictive structure that are not reflected by conventional forecast errors. The framework thereby turns domain-expert representations into systematic tests of learned physical dynamics without prescribing the relationships in advance.
Oskar Bohn Lassen, Joao Paulo de Souza Boger, Simon Driscoll +4
Department of Technology, Management, and Economics Technical University of Denmark · Department of Applied Mathematics and Theoretical Physics University of Cambridge · Department of Mathematics and Statistics University of Exeter
Large language models (LLMs) show strong capabilities in general reasoning but typically lack reliability in scientific domains like quantum mechanics, which demand strict adherence to physical constraints. This limitation arises from the scarcity of verifiable training resources and the inadequacy of coarse feedback signals in standard alignment paradigms. To address the data challenge, we introduce QuantumQA, a large-scale dataset constructed via a task-adaptive strategy and a hybrid verification protocol that combines deterministic solvers with semantic auditing to guarantee scientific rigor. Building on this foundation, we propose the verification-aware reward model (VRM) tailored for Reinforcement Learning with Verifiable Rewards (RLVR), which employs an adaptive reward fusion (ARF) mechanism to dynamically integrate deterministic signals from a scientific execution suite (SES) with multidimensional semantic evaluations for precise supervision. Experimental results demonstrate that our method consistently outperforms baselines and general-purpose preference models. Notably, our optimized 8B model achieves performance competitive with proprietary models, validating that incorporating verifiable, rule-based feedback into the reinforcement learning loop offers a parameter-efficient alternative to pure scaling.
Songxin Qu, Tai-Ping Sun, Yun-Jie Wang +8
Institute of Advanced Technology, University of Science and Technology of China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center · School of Physics, University of Science and Technology of China +2