Evaluating Dynamical Fidelity through Predictive Structure in Physical Representations
Authors: Oskar Bohn Lassen, Joao Paulo de Souza Boger, Simon Driscoll, Stephen I. Thomson, Sebastian Schemm, Filipe Rodrigues, Francisco C. Pereira
Organizations: Department of Technology, Management, and Economics Technical University of Denmark · Department of Applied Mathematics and Theoretical Physics University of Cambridge · Department of Mathematics and Statistics University of Exeter
Machine-learning models for physical systems are currently evaluated primarily through errors between predicted and reference states and, increasingly, through tests of physical consistency. These metrics assess whether predictions are accurate and satisfy selected physical requirements, but provide limited insight into whether learned trajectories reproduce the underlying dynamics. Domain experts examine such relationships through physical representations that expose relevant processes, interactions, and responses, but these analyses are often separated from typical machine-learning evaluation. We introduce a practical framework for evaluating dynamical fidelity through predictive structure in physical representation spaces. Experts define the representations, while reference trajectories determine which relationships are predictive and retained as evaluation tests. We demonstrate the approach in atmospheric forecasting using ERA5 representations of planetary-wave activity and Northern Annular Mode evolution, and evaluate Pangu-Weather, GraphCast, and FengWu. The models exhibit distinct departures from reference predictive structure that are not reflected by conventional forecast errors. The framework thereby turns domain-expert representations into systematic tests of learned physical dynamics without prescribing the relationships in advance.
Figures & tables
Figure 1: First-level predictive structure in ERA5 before robustness and permutation-FDR filtering. (a) Maximum FVE across the candidate space by latitude–pressure region. (b) Representation/wavenumber deep-dive for the three highest-FVE regions. (c) Window/operator deep-dive for their strongest configurations. Outlined cells indicate the successive maxima.
Figure 2: Selection of robust predictive relationships. (a) Of 7,875 first-level candidates, 295 satisfy the 95% directional-stability criterion; permutation-FDR calibration retains 283 at FVE ≥0.0504 . (b) Second-level search within their frozen branches evaluates 4,358,681 conditional candidates, of which 71,494 satisfy 95% stability and 672 remain after permutation-FDR calibration (conditional FVE ≥0.3102 ). (c) Example of leaf that deviates from the neutral mean, isolating initially neutral states that evolve strongly. Lines show winter-weighted medians and shading 10–90% ranges.
Figure 3: Model evaluation across the 672 selected terminal conditions associated with the retained second-level relationships. (a–b) Absolute error in the conditional mean and variance of future 50–100 hPa NAM within each frozen ERA5-defined terminal leaf; lines show the mean across relationships and shading ±1 standard deviation across relationships. (c) Conventional T , u , and v forecast MAE on the same subsets (unweighted global mean over the 1∘ grid and 13 pressure levels).
GraphCast
FengWu
Pangu-Weather
Fidelity
Forecast
Fidelity
Forecast
Fidelity
Forecast
Predictive structure
N
%
Eμ
Eσ2
T
u
v
Eμ
Eσ2
T
u
v
Eμ
Eσ2
T
u
v
All
672
100.0
0.17
0.20
2.29
5.73
5.77
0.15
0.21
1.76
4.49
4.56
0.09
0.11
2.17
5.49
5.58
Representations
Total EP1
102
15.2
0.21
0.18
2.30
5.74
5.76
0.18
0.20
1.77
4.50
4.57
0.09
0.10
2.19
5.53
5.61
Wave-1 EP1
96
14.3
0.15
0.26
2.29
5.71
5.77
0.11
0.21
1.77
4.50
4.57
0.09
0.11
2.16
5.45
5.55
Wave-2 EP1
12
1.8
0.26
0.20
2.27
5.68
5.69
0.18
0.17
1.75
4.45
4.52
0.11
0.07
2.17
5.47
5.55
Table 1: Model evaluation over lead days 6–10 for selected subsets of the 672 selected terminal conditions. Rows are defined by representations used in at least one split and may therefore overlap. Eμ and Eσ2 denote conditional NAM mean and variance errors; T , u , and v are conventional forecast MAE (unweighted global mean over the 1∘ grid and 13 pressure levels). Dark green and light green indicate the lowest and second-lowest model errors in each column, respectively.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Schematic illustration of the zonal-mean decomposition and zonal eddy covariances. At a fixed latitude, pressure, and time, each atmospheric variable forms a longitude-dependent field around an east–west circle. Dashed lines show the zonal means, while the shaded departures are the eddies u′ , v′ , and T′ . Multiplying the relevant departures longitude by longitude and averaging the products around the circle gives u′v′ and v′T′ . The faded globes indicate that the same calculation is repeated independently across latitude and pressure. The calculation is also repeated for every time-step but this is not included in the illustration above.
Model
Average over
Eμ
Eσ2
T
u
v
GraphCast
672 relationships
0.173
0.195
2.290
5.729
5.768
287 motifs
0.171
0.179
2.285
5.717
5.748
FengWu
672 relationships
0.148
0.210
1.762
4.488
4.558
287 motifs
0.169
0.205
1.765
4.489
4.556
Pangu-Weather
672 relationships
0.091
0.107
2.175
5.494
5.581
287 motifs
0.090
0.104
2.180
5.498
5.579
Appendix
Table 2: Lead-day 6–10 errors averaged over the 672 retained relationships and over the 287 motifs (relationships averaged within each motif, motifs weighted equally). T , u , and v are forecast MAE in K and m s -1 .
Machine learning weather prediction (MLWP) models have achieved impressive forecasting performance at a small fraction of the computational costs required for traditional physics-based methods. However, they are primarily (1) data-driven and (2) evaluated using pixel-wide error metrics (e.g., RMSE), so there are no guarantees that their forecasts are consistent with known physical laws. We introduce PhysMetrics.Weather, an evaluation framework that assesses the physical realism of MLWP models across three types of metrics: conservation, spectral, and dynamical. By quantifying physical realism, this tool guides the development of physics-informed architectures and helps evaluate whether MLWP models are reliable for operational use. Our framework is available on Github at https://github.com/Emmakast/PhysMetrics.Weather.
Emma Kasteleyn, Timo Maier, Axel Lauer +3
University of Amsterdam, Amsterdam, The Netherlands · Deutsches Zentrum für Luft- und Raumfahrt (DLR), Institut für Physik der Atmosphäre, Oberpfaffenhofen, Germany · University of Bremen, Institute of Environmental Physics (IUP), Bremen, Germany +1
Multivariate forecasting in physical systems requires models that predict coupled temporal variables while preserving meaningful state evolution. Deep forecasters can fit temporal correlations, and physics-informed models can regularize predictions with scientific constraints, but these directions are often connected only at the decoded-output level. As a result, the hidden predictive state that generates future trajectories may remain statistically useful but physically unstructured. We introduce Phys-JEPA, a physics-informed joint-embedding predictive architecture for multivariate time-series forecasting. Phys-JEPA learns a latent world model in which predictive states are decomposed into physical and residual components, and physical consistency is imposed directly on latent states and latent transitions rather than only on decoded forecasts. This formulation uses known physical variables to organize the representation space while retaining residual capacity for unresolved dynamics. On Jena Climate 2009--2016, Phys-JEPA reduces aggregate MSE from 0.12482 to 0.12273 and temperature MSE from 0.01892 to 0.01831 at H=24. On Traffic, full Phys-JEPA improves aggregate MSE over the supervised baseline across all tested horizons, reducing H=192 MSE from 0.800784 to 0.773873. On Electricity, the best variant depends on horizon: static latent consistency is strongest at H=24 and H=48, while full Phys-JEPA gives the best aggregate and target-variable MSE at H=192. These initial results suggest that moving physics-informed learning from output space to latent predictive state space is a promising direction for interpretable temporal world models.
Despite their high accuracy on point-wise metrics, machine learning weather forecasting models can exhibit different failure modes such as blurring, periodic irregularities, and other unphysical spatial artifacts. This has motivated a variety of metrics to detect known failure cases. Existing metrics fix a representation or transformation in advance, and that choice limits the artifacts they can detect. We propose to train a discriminator for separating reference data from the model's output, and using its output logit to obtain a divergence-like realism score. The discriminator learns whatever separates the model's fields from real weather, adapting to whichever failure mode that model exhibits. We compare our learned atmospheric critic to existing metrics using various synthetic corruptions applied to ERA5 reanalysis data. Our method successfully identifies the corruptions and ranks their severity, while existing metrics fail on at least one corruption. Additionally, we evaluate forecasts from real weather models, and find that the realism score degrades with longer lead times and the metric generally assigns higher realism to numerical models than to machine learning models.
Younes Elberkennou, Dmitri Demler, Thierry Meier +3
ETH Zürich, Switzerland · ETH AI Center, ETH Zürich, Switzerland