Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices actually improve real-world data assimilation. We present the first controlled benchmark of generative weather data assimilation on real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four weather variables, we evaluate methods while holding the dataset, observation operator, and deep learning architecture fixed. The benchmark compares the major design choices, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three clear conclusions. First, learned generative priors outperform the Gaussian prior of 3D-Var (35.7% vs. 33.3% RMSE reduction over ERA5) despite using no ERA5 background field at inference. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. We further evaluate both dense and sparse station settings and find advantages from generative AI and full-gradient guidance more pronounced under sparsity. Together, these results identify which components of generative weather data assimilation improve performance on real station observations and establish a standardized benchmark for future work.
Figures & tables
Figure 1 : Conceptual illustration of generative data assimilation: starting from sparse station observations on the four target variables, a generative prior trained on ERA5 produces a full gridded analysis consistent with the observations.
Diffusion Model
Flow Matching
Network output
Denoiser Dθ≈x1
Velocity vθ≈x1−x0
Training schedule
lnσ∼N(Pmean,Pstd2)
t∼U(0,1)
Score function
(sDθ−x)/(s2σ2)
(tvθ−x)/(1−t)
Sampler
EI predictor + Langevin corrector
Adaptive ODE / Euler / Euler + corrector
Table 1 : Comparison of diffusion model and flow matching. Pmean and Pstd parameterize the EDM noise-sampling distribution ( Appendix B ).
Figure 2 : Overview of the data assimilation strategies compared in this work. (a) Classical methods : 3D-Var and its latent variant minimize a background-plus-observation cost with no generative prior. (b) Diffusion-SDA ( Section 3.3.1 ): per-step guidance through the denoiser Dθ , optionally followed by a Langevin corrector, toward the observation-consistent region H(x)=y∗ . (c) Flow Guidance ( Section 3.3.2 ): the flow-matching analogue, backpropagating the observation loss through the terminal extrapolation x^1(xt) . (d) FlowDPS [ 28 ] : iterates K optimization steps on the terminal estimate x^1(K)(xt) and resamples to the next ODE step. (e) Latent Flow Guidance : Flow Guidance inside the latent space of an autoencoder E/D ; the observation loss is evaluated in pixel space via the decoder. (f) D-Flow ( Section 3.3.3 ): no per-step guidance; an outer loop updates the initial noise x0(0)→x0(k) by backpropagating through the full ODE until x1 satisfies the observations. Numbered circles on panel (c) mark the four design axes ablated in Appendix C : (1) guidance schedule λ(t) , (2) stop gradient, (3) terminal extrapolation x^1 , (4) sampler; a fifth axis (pixel vs. latent) is captured by the (c)/(e) contrast.
Method
Model class
Space
DA mechanism
M
3D-Var
Classical
Pixel
Variational optimization
1
Latent 3D-Var
Classical
Latent
Variational optimization
1
Diffusion-SDA
Diffusion
Pixel
Full-gradient guidance
16
Flow Guidance
Flow matching
Pixel
Full-gradient guidance
16
FlowDPS
Flow matching
Pixel
Stop-gradient guidance
16
D-Flow
Flow matching
Pixel
Initial-noise optimization
1
Table 2 : Summary of data assimilation methods. Columns: model class (classical, diffusion, or flow matching), operating space (pixel or learned latent), DA mechanism (how observations are injected), and default ensemble size M .
Figure 3 : Dataset split and regional station density under the two benchmarks. Top block (Dense, 11,849 stations): U.S. map with stations colored by partition (Train: 8,294 / Val: 1,777 / Test: 1,778) and dashed boxes marking three zoom regions, followed by California, Midwest, and New England panels in Lambert Conformal projection with identical physical extent ( 800×800 km). Bottom block (Sparse, 1,000 stations): same layout for a uniformly subsampled set (Train: 700 / Val: 150 / Test: 150). Matching extents enable fair density comparison across regions and between the dense and sparse benchmarks.
u10
v10
T2
Td,2
Average
Cost
Method
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Mem
Total
3D-Var
51.7
35.2
51.8
35.1
46.0
29.2
49.5
33.6
49.7
33.3
0.17
4
3D-Var (w/o ERA5 loss)
66.0
15.0
66.2
13.7
61.8
9.4
64.8
12.3
64.7
12.6
0.20
36
Diffusion-SDA
52.7
37.8
52.2
38.6
39.9
29.6
47.6
35.6
48.1
35.4
6.12
246
Diffusion-SDA (PO)
48.3
37.9
47.9
38.3
29.5
19.9
36.7
25.0
40.6
30.3
6.11
82
D-Flow
50.7
34.7
49.6
35.8
33.3
23.0
39.8
28.5
43.4
30.5
36.01
2563
Table 3 : Dense benchmark on the 11,849 -station network. RMSE improvement (%) over ERA5 at training (Tr) and held-out test (Te) stations; “Average” is the unweighted mean across the four variables. “Mem” is peak GPU memory (GB); “Total” is wall-clock time (min) for 360 snapshots on a single RTX8000. Higher improvement and lower cost are better. “PO” = predictor-only (no Langevin corrector).
u10
v10
T2
Td,2
Average
Cost
Method
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Mem
Total
3D-Var
49.0
14.1
49.9
15.4
46.6
11.7
48.8
14.1
48.6
13.8
0.11
3
Latent 3D-Var ( S=2 , G=1 )
71.2
14.6
71.5
16.4
65.7
13.3
69.2
15.5
69.4
15.0
0.25
11
Latent 3D-Var ( S=2 , G=4 )
68.1
15.2
68.8
16.4
56.6
13.6
63.2
15.5
64.2
15.2
0.34
13
Diffusion-SDA
69.3
25.7
69.0
28.1
46.6
14.2
60.0
17.9
61.2
21.5
6.12
242
Diffusion-SDA (PO)
62.0
28.6
61.1
28.6
30.3
−2.4
42.6
−2.8
49.0
13.0
6.10
83
Table 4 : Sparse benchmark, 1,000 randomly subsampled stations. Columns as in Table 3 . “SG” applies the stop-gradient approximation; “Euler-128” replaces the adaptive dopri5 solver with a 128 -step fixed-step Euler solver; “PO” = predictor-only; “PC” adds a Langevin corrector.
Figure 4 : Macro RMSE improvement ( Δ% ) over ERA5 for Flow Guidance, Diffusion-SDA, and 3D-Var, as a function of mean 5-NN distance to assimilated stations, for (a) the dense and (b) the sparse benchmark. Lines show the median across 360 time steps; shaded bands span the 10th–90th percentile range. Vertical dashed lines mark the median 5-NN distance of a random CONUS location under each benchmark ( 48 km dense; 142 km sparse). In the sparse benchmark, beyond ∼150 km temperature and dewpoint improvements from the generative methods turn negative while wind improvements remain positive.
Figure 5 : Rocky Mountains case study for T2 . Top row: geographic context (left) and satellite view (right) of the region of interest. Middle row: (a) ERA5 background, (b) Flow Guidance, (c) Diffusion-SDA, (d) 3D-Var. Bottom row: per-station errors (analysis − observation) for each method, with mean absolute error (MAE) and RMSE annotated. ERA5 shows a near-uniform ∼6.5 K cold bias across the region; all three methods substantially reduce the error, with the generative methods (RMSE 3.3 – 3.5 K) modestly outperforming 3D-Var (RMSE 4.0 K).
Figure 6 : Great Lakes case study for u10 . Panels as in Figure 5 . ERA5 overestimates wind speed by ∼4.5 m/s across the region. Flow Guidance and Diffusion-SDA reduce the error at the shoreline test stations (RMSE 2.39 m/s, 53.0% reduction), while 3D-Var retains larger residual errors (RMSE 2.81 m/s, 44.8% reduction).
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Interpolation
u10
v10
T2
Td,2
Average
ERA5 Baseline
Nearest
1.90 ± 0.31
1.95 ± 0.28
2.51 ± 0.31
2.72 ± 0.51
2.27 ± 0.35
Bilinear
1.88 ± 0.31
1.93 ± 0.28
2.43 ± 0.30
2.60 ± 0.45
2.21 ± 0.34
Bicubic
1.89 ± 0.31
1.94 ± 0.28
2.44 ± 0.31
2.63 ± 0.47
2.22 ± 0.34
Flow Guidance
Bilinear
1.15 ± 0.17
1.18 ± 0.18
1.75 ± 0.18
1.70 ± 0.21
1.44 ± 0.18
Appendix
Table 5: Effect of interpolation scheme on ERA5 baseline and Flow Guidance assimilation RMSE (mean ± std, validation stations). Bicubic is selected based on validation performance. Bold marks the best value per column within each panel.
u10
v10
T2
Td,2
Training std (physical units, m s -1 or K)
3.45
3.84
11.70
11.67
ERA5 validation RMSE (physical)
1.89
1.94
2.44
2.63
ERA5 validation RMSE (normalized)
0.547
0.504
0.209
0.225
Appendix
Table 6 : ERA5 baseline error in physical and normalized units.
Variant
Covariance Σ
Notes
DPS [ 9 ]
Σy
Ignores uncertainty in x^1 .
SDA / Π GDM [ 47 ]
Σy+HΣ0(t)H⊤
Σ0=γσ2I (SDA) or rt2I ( Π GDM).
MMPS [ 2 ]
Σy+HV[x1∣xt]H⊤
Exact posterior covariance; conjugate-gradient solves per step.
NDTM [ 39 ]
per-step variational optimization
Most general; high per-step cost.
Appendix
Table 7 : Common Gaussian likelihood approximations p(y∗∣xt)≈N(y∗;H(x^1),Σ) used in posterior-guidance methods. We use the SDA form throughout.
Table 8 : Ablation: guidance schedule. Top : effect of removing guidance from 20%-wide windows (constant schedule, λ=2×105 ). Bottom : comparison of four schedule shapes ( Figure 7 ), each at its optimal λ . All use dopri5, single-step terminal extrapolation, without stop gradient. RMSE improvement (%) over ERA5 at training (Tr) and test (Te) stations; total wall-clock cost (min) for 360 snapshots.
Figure 7 : The four guidance schedules compared in Table 8 . The constant schedule maintains uniform guidance strength throughout the trajectory; cosine-increasing concentrates guidance near t=1 ; cosine-decreasing concentrates it near t=0 ; and the SDA-derived schedule ( Equation 31 ) peaks at intermediate t and vanishes at both endpoints.
u10
v10
T2
Td,2
Avg
Cost
Method
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Mem
Total
w/o SG (default)
53.2
38.2
53.0
38.8
42.9
29.9
49.1
35.9
49.6
35.7
3.05
380
w/ SG
60.9
34.1
61.0
34.9
54.6
27.8
58.1
32.0
58.6
32.2
0.37
77
FlowDPS (terminal-space optimization)
FlowDPS, K=5
52.7
35.5
52.5
35.7
27.3
2.1
38.2
12.0
42.7
21.3
0.25
28
FlowDPS, K=10
57.8
34.5
57.7
35.2
45.2
16.9
51.5
22.6
53.0
27.3
0.25
31
Appendix
Table 9 : Ablation: stop gradient (SG) and FlowDPS. RMSE improvement (%) over ERA5 at training (Tr) and test (Te) stations, and total wall-clock cost (min) for 360 snapshots. Without SG, gradients backpropagate through the velocity network. FlowDPS- K denotes FlowDPS with K optimization steps per ODE step.
u10
v10
T2
Td,2
Avg
Cost
Extrapolation
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Mem
Total
Single-step (default)
53.2
38.2
53.0
38.8
42.9
29.9
49.1
35.9
49.6
35.7
3.05
380
RK4 two-step
55.7
37.6
55.4
38.3
46.4
31.0
51.7
36.4
52.3
35.8
25.21
10,659
Appendix
Table 10 : Ablation: terminal extrapolation. RMSE improvement (%) over ERA5 at training (Tr) and test (Te) stations, plus total wall-clock cost (min) for 360 snapshots.
u10
v10
T2
Td,2
Avg
Cost
Sampler
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Tr
Te
Mem
Total
Flow Guidance dopri5 (PO, default)
53.2
38.2
53.0
38.8
42.9
29.9
49.1
35.9
49.6
35.7
3.05
380
Flow Guidance Euler-128 (PO)
49.2
38.6
48.8
39.0
34.7
24.1
40.8
29.6
43.4
32.8
2.92
53
Flow Guidance Euler-128 (PC)
48.9
38.7
48.7
39.3
35.4
26.9
42.9
34.5
44.0
34.9
2.93
208
Diffusion-SDA PC (default)
52.7
37.8
52.2
38.6
39.9
29.6
47.6
35.6
48.1
35.4
6.12
246
Diffusion-SDA PO
48.3
37.9
47.9
38.3
29.5
19.9
36.7
25.0
40.6
30.3
6.11
82
Appendix
Table 11 : Ablation: ODE solver and corrector. RMSE improvement (%) over ERA5 at training (Tr) and test (Te) stations, plus total wall-clock cost (min) for 360 snapshots. Flow Guidance dopri5 uses λ=2×105 ; Euler-128 variants use λ=105 . “PO” = predictor-only; “PC” adds a Langevin corrector between predictor steps.
Figure 8 : Perturbation amplification ( log10 scale) as a function of injection time t , averaged over 16 trajectories with δ=0.01 . Top row (a,b): Isolate score quality by fixing the flow ODE trajectory. Flow PO/PC use the flow-derived score ( τ=0.02 ); Hybrid PO/PC use the diffusion denoiser’s score ( τ=0.1 ) on the same trajectory. Bottom row (c,d): Isolate trajectory by fixing the diffusion score. Hybrid PO/PC integrate the flow ODE; Diffusion PO/PC integrate the VP-SDE ( τ=0.1 ). Left panels show guidance-direction perturbations; right panels show random-direction perturbations. Insets zoom into t∈[0.95,1.0] .
Figure 9 : Accuracy–compute trade-off for Flow Guidance in pixel vs. latent space. x -axis: total wall-clock cost for 360 snapshots on 1 GPU (log scale). y -axis: test-set RMSE improvement over ERA5, averaged over {u10,v10,T2,Td,2} . Colors denote downsampling depth S ; marker shapes denote the attention grouping (circle: G=1 , variable-plus-spatial-mixing; square: G=4 , spatial-mixing-only). Lines connect configurations of the same architecture across guidance strengths λ . Peak memory (approximately constant within each S -group) is annotated per group. Reported runtimes for G=4 at S≥3 are inflated by a PyTorch grouped-convolution inefficiency that a mathematically-equivalent batch-reshape implementation removes ( 1.5 – 4× faster).
3D-Var variant
Tr (%)
Te (%)
Pixel (no AE)
49.7
33.3
Latent ( S=2 , G=1 )
53.4
33.3
Latent ( S=2 , G=4 )
55.5
33.5
Appendix
Table 12 : 3D-Var accuracy in three nonlinear regimes. RMSE improvement (%) over ERA5 at training (Tr) and test (Te) stations, averaged across the four variables. The three variants are indistinguishable on Te.
G
S
Compr.
u10
v10
T2
Td,2
MAE
4
2
0.99 ×
0.0069
0.0078
0.0175
0.0177
0.0125
4
3
3.85 ×
0.0819
0.0808
0.1482
0.1735
0.1211
4
4
14.7 ×
0.1982
0.2024
0.2688
0.3322
0.2504
1
2
0.99 ×
0.0063
0.0067
0.0213
0.0218
0.0140
1
3
3.85 ×
0.0533
0.0573
0.1494
0.1897
0.1124
1
4
14.7 ×
0.1307
0.1377
0.3185
0.3720
0.2397
Appendix
Table 13 : Autoencoder reconstruction MAE on test fields in normalized units. “Compr.” is the spatial compression factor.
Figure 10 : Per-variable RMSE improvement vs. mean 5-NN distance, dense benchmark.
Figure 11 : Per-variable RMSE improvement vs. mean 5-NN distance, sparse benchmark.
Figure 12 : Los Angeles basin case study for T2 . Top row: geographic context (left) and satellite view (right) of the region of interest. Middle row: (a) ERA5 background, (b) Flow Guidance, (c) Diffusion-SDA, (d) 3D-Var. Bottom row: per-station errors (analysis − observation) for each method, with MAE and RMSE annotated. ERA5 overestimates mountain temperatures and underestimates coastal urban temperatures; all three methods achieve nearly identical corrections on this case (RMSE 3.6 – 3.7 K, 36 – 37% reduction).
Figure 13 : Florida peninsula case study for u10 . Panels as in Figure 12 . Flow Guidance and Diffusion-SDA reduce wind RMSE by ∼72% on this snapshot versus 66% for 3D-Var, a ∼6 pp gap that mirrors the Great Lakes finding at smaller magnitude.
u10
v10
T2
Td,2
Average
Method
Tr
Va
Te
Tr
Va
Te
Tr
Va
Te
Tr
Va
Te
Tr
Va
Te
ERA5 (baseline)
1.87
1.89
1.88
1.93
1.94
1.91
2.40
2.44
2.38
2.62
2.63
2.53
2.20
2.22
2.18
3D-Var
0.90
1.21
1.21
0.93
1.26
1.24
1.29
1.76
1.68
1.31
1.76
1.66
1.11
1.50
1.45
Diffusion-SDA
0.88
1.16
1.16
0.92
1.18
1.17
1.43
1.73
1.67
1.36
1.69
1.61
1.15
1.44
1.40
Diffusion-SDA (PO)
0.96
1.16
1.16
1.00
1.19
1.18
1.68
1.93
1.90
1.63
1.95
1.87
1.32
1.56
1.52
D-Flow
0.92
1.21
1.22
0.97
1.23
1.22
1.59
1.87
1.82
1.55
1.84
1.79
1.26
1.54
1.51
Appendix
Table 14 : Absolute RMSE (train/validation/test) on the dense benchmark. Wind variables are in m s -1 ; temperature and dewpoint are in K.
u10
v10
T2
Td,2
Average
Method
Tr
Va
Te
Tr
Va
Te
Tr
Va
Te
Tr
Va
Te
Tr
Va
Te
Absolute RMSE
ERA5
1.86
1.87
1.87
1.93
1.94
1.92
2.40
2.44
2.39
2.61
2.62
2.52
2.20
2.22
2.17
3D-Var
0.90
1.21
1.21
0.93
1.26
1.23
1.29
1.76
1.69
1.31
1.76
1.66
1.11
1.50
1.45
Diffusion-SDA
0.87
1.15
1.16
0.92
1.18
1.17
1.43
1.73
1.67
1.36
1.70
1.61
1.14
1.44
1.40
Flow Guidance
0.87
1.14
1.15
0.90
1.17
1.16
1.36
1.73
1.67
1.32
1.69
1.60
1.11
1.43
1.40
Appendix
Table 15 : Full-year benchmark on the dense station network ( 8,664 snapshots, 2023). RMSE in physical units (top) and improvement over ERA5 (bottom).
Information Sciences and Technology The Pennsylvania State University University Park, Pennsylvania · School of Mechanical Engineering Purdue University West Lafayette, Indiana
School of Mechanical Engineering, Purdue University, West Lafayette, Indiana, USA · College of Information Sciences and Technology, The Pennsylvania State University, University Park, Pennsylvania, USA