ProtocolMatch: Protocol-Dependent Model Selection for Scientific Dynamics Forecasting
Organizations: Stony Brook University · Westlake University
Abstract
Scientific dynamics forecasting is often framed as an architecture choice, although deployment is also determined by observed history, rollout feedback, compute budget, physical objective, and test distribution. We formulate protocol-dependent model selection and introduce ProtocolMatch, a compute-matched, validation-selected, and failure-preserving evaluation framework. On driven quantum-spin dynamics, we compare recurrent, patched-attention, causal-attention, and low-rank linear predictors across three independently generated datasets. The causal-attention--recurrence ordering reverses as the training set grows within a fixed two-spin task, while a linear predictor has the lowest mean error in the six-spin local-observable comparison. Restricting observed history worsens every refreshed-history view but improves every closed-loop view in the four-spin study. A latest-state MLP has lower error than persistence on every dataset under state refresh across all five cells, yet its closed-loop rank varies by system and includes finite explosive errors. Physical penalties improve targeted consistency without reliably improving prediction error, and in-distribution intervals lose most coverage after a driving-frequency shift. Thus scientific model selection should return a predictor with its protocol and report accuracy, physical validity, and shifted-distribution reliability separately.
Figures & tables
| Factor | Values | Question |
|---|---|---|
| History | 1, 5, 20 | How much past state is useful? |
| Channels | Full, local | Which observables are available? |
| Rollout | Refresh, closed loop | Is feedback autonomous? |
| Clipping | Off, on | Are bounds imposed at inference? |
| Budget | 30 s, 120 s | Is optimization time matched? |
| Test family | ID, freq., state shifts | Where must the model generalize? |
| System and target | Train traj. | LSTM | PatchTST | Causal Transformer | Low-rank linear |
|---|---|---|---|---|---|
| Two-spin Ising, full Pauli | 64 | 0.09402 | 0.65869 | 0.08681 | 0.19583 |
| Two-spin Ising, full Pauli | 256 | 0.02456 | 0.66745 | 0.04654 | 0.18849 |
| Four-spin Ising, full Pauli | 256 | 0.10608 | 0.12416 | 0.10902 | 0.12220 |
| Two-spin XXZ, full Pauli | 256 | 0.03074 | 0.25047 | 0.04540 | 0.18421 |
| Six-spin Ising, local Pauli | 256 | 0.08833 | 0.08835 | 0.07878 | 0.07668 |
| Observed-history refresh | Unclipped closed loop | ||||
|---|---|---|---|---|---|
| System and target | Train traj. | Current-state MLP | Persistence | Current-state MLP | Persistence / direction |
| Two-spin Ising, full Pauli | 64 | / mixed | |||
| Two-spin Ising, full Pauli | 256 | / MLP 3/3 | |||
| Four-spin Ising, full Pauli | 256 | / mixed | |||
| Two-spin XXZ, full Pauli | 256 | / persistence 3/3 | |||
| Six-spin Ising, local Pauli | 256 | / MLP 3/3 | |||
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | Evaluations | Independent datasets | Primary purpose |
| Core matched comparisons | 288 | 3 per system cell | Architecture and data-regime ranks |
| History and input controls | 540 | 3 per condition | Information–rollout interaction |
| Constraint terms and weights | 216 | 1 | Accuracy–validity trade-off |
| Six-spin extension | 72 | 3 | Local-observable scaling check |
| Primary matrix total | 1,116 | — | — |
| Diagnostic baselines (separate) | 105 | 3 per system cell | Latest-state and persistence controls |
| Contrast | Condition | Dataset 1 | Dataset 2 | Dataset 3 | Mean |
|---|---|---|---|---|---|
| LSTM minus causal | 64 trajectories | +0.01292 | +0.00508 | +0.00363 | +0.00721 |
| LSTM minus causal | 256 trajectories | -0.02696 | -0.01586 | -0.02314 | -0.02198 |
| Local minus full input | Observed-history | +0.01967 | +0.02011 | +0.02162 | +0.02046 |
| Local minus full input | Closed-loop | -0.03221 | -0.03072 | -0.02776 | -0.03023 |
| System | Comparison | Lower-error views | All-data agreement |
|---|---|---|---|
| Two-spin Ising | LSTM history 5 vs. 20 | 27/32 | 9/32 |
| Four-spin Ising | LSTM history 5 vs. 20 | 30/32 | 28/32 |
| Four-spin Ising | Local vs. full, observed history | 0/16 | 16/16 worse |
| Four-spin Ising | Local vs. full, closed loop | 16/16 | 16/16 better |
| Penalty weight | Views with lower trace error | Views with higher prediction MSE |
| 0.001 | 31/32 | 26/32 |
| 0.01 | 32/32 | 17/32 |
| 0.1 | 32/32 | 21/32 |
| Predictor | ID coverage | Shift coverage | Width |
|---|---|---|---|
| LSTM | 87.50% | 28.13% | 1.2767 |
| PatchTST | 91.32% | 26.91% | 7.6545 |
| Causal Transformer | 86.98% | 16.49% | 1.7195 |
| Low-rank linear | 87.67% | 24.31% | 2.3278 |