Station weather forecasting is fundamentally shaped by both complex spatial dependencies across stations and strong physical coupling among weather variables. However, existing studies often consider these relationships separately and use different datasets and experimental settings, hindering systematic assessment of their individual and joint contributions. In this paper, we introduce M2Weather, a benchmark for joint multi-station and multi-variable weather forecasting. Through multi-criteria quality control and station stratification, we collect 2,809 high-quality stations with 5 physically coupled weather variables across three spatial scales: France, Europe, and Global. This multi-scale design lets us examine whether conclusions persist from national to global station networks. We also introduce unified training and evaluation protocols to enable fair comparison of different station-variable modeling paradigms. To further examine the benefits of modeling station-variable relationships, we design a lightweight, plug-and-play adapter. With a trained weather forecasting model, this adapter can introduce missing station or variable relationships without retraining the model. This enables fair and efficient investigation of station-variable relationships. Systematic evaluation of 16 representative models shows the benefits of jointly modeling station and variable relationships. Completing missing relationships further reduces MSE for all adapted models on all three datasets. Together, these results identify the complementary information across stations and variables as an important resource for improving station weather forecasting. Our code can be obtained at https://github.com/hnu-vis/M2-Weather.
Figures & tables
Figure 1: Four station–variable forecasting paradigms: (a) SSSV models each series independently; (b) SSMV captures interactions among variables within a station; (c) MSSV captures interactions across stations for one variable; and (d) MSMV models both. All four use temporal history, but differ in the station and variable context available to each forecast.
Figure 2: Overview of the M 2 Weather benchmark and interaction analysis framework. (a) Dataset construction and station stratification based on temporal alignment, quality, provenance, and metadata. (b) Unified protocol for heterogeneous backbones, with model-specific wrappers preserving native interaction structures. (c) Adapter-based analysis on frozen benchmark backbones, with calibration and variable/station interaction corrections fitted by gradient-free ridge regression.
Figure 3
Paradigm
Baselines
SSSV
DLinear ( Zeng et al., 2023 ) , xPatch ( Stitsyuk and Choi, 2025 ) , PatchTST ( Nie et al., 2023 ) , Timer ( Liu et al., 2024b ) , TimeMoE ( Shi et al., 2025 )
SSMV
iTransformer ( Liu et al., 2024a ) , TQNet ( Lin et al., 2025 ) , DUET ( Qiu et al., 2025 ) , Moirai ( Woo et al., 2024 ) , Timer-XL ( Liu et al., 2025 )
MSSV
Corrformer ( Wu et al., 2023 ) , S 2 Transformer ( Chen et al., 2026 ) , CDPNet ( Xu et al., 2025b ) , EasyST ( Tang et al., 2024 ) , STELLA ( Fu et al., 2025 )
MSMV
HiSTGNN ( Ma et al., 2023 )
Table 2: Baseline coverage under the four station–variable modeling paradigms in M 2 Weather.
France
Europe
Global
Paradigm
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
SSSV
DLinear
0.4977
0.4605
−7.49
0.4481
0.4087
−8.79
0.4508
0.4172
−7.45
xPatch
0.3974
0.3642
−8.36
0.3509
0.3280
−6.53
0.3587
0.3371
−6.04
PatchTST
0.3905
0.3601
−7.80
0.3472
0.3266
−5.92
0.3540
0.3341
−5.61
Timer
0.5211
0.4445
−14.70
0.4392
0.3845
−12.47
0.4547
0.3986
−12.33
TimeMoE
0.4354
0.3845
−11.69
0.3798
0.3435
−9.56
0.3943
0.3559
−9.73
Table 3: MSE ( 48→72 hours). Δ is the relative change after adaptation; negative values indicate lower error. HiSTGNN is reported without an adapter. Underlining marks the best native backbone; bold marks the best displayed result (ties at four decimals). Per-variable MSE, MAE, and standard deviations are in Appendix C.3 .
Figure 5: MSE by modeling paradigm, averaged over models. Connected markers compare native and adapted predictions; percentages denote reductions in group means. Parentheses indicate group sizes. HiSTGNN is the native MSMV reference, without an adapter.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Criterion
M 2 -CS
M 2 -RA
Five-year coverage of every variable
≥0.82
≥0.20
Five-year mean coverage
–
≥0.65
Monthly minimum coverage, P10
≥0.60
–
Monthly fourth-highest coverage, P20
–
≥0.50
Monthly mean coverage, P20
–
≥0.50
Monthly maximum five-variable gap, P90
≤168 h
–
Appendix
Table 4: Station-tiering criteria. A dash indicates that the criterion is not used for the corresponding tier.
Code
Source
Present in release
0
Unresolved missing value
No
1
Quality-controlled exact-hour observation
Yes
2
Quality-controlled observation within 15 minutes
Yes
3
Same-station interpolation of a gap ≤12 h
Yes
4
Nearest ERA5 grid cell at the same hour
Yes
5
Same ERA5 grid cell at a nearby hour
No
Appendix
Table 5: Per-value source codes stored in the imputation mask.
Dataset
T
Td
WS
P
WD
Overall
Global
1.392
1.953
1.578
2.343
5.576
2.568
Europe
2.608
3.216
3.630
11.242
7.499
5.639
France
0.585
0.650
0.735
1.266
1.381
0.923
Appendix
Table 6: Percentage of values completed by either station interpolation or ERA5.
Split
Start
End
Valid windows
Training
2017-01-01 00:00
2020-01-01 14:00
26,175
Validation
2020-01-01 14:00
2020-07-02 04:00
4,263
Test
2020-07-02 04:00
2022-01-01 00:00
13,029
Appendix
Table 7: Shared temporal splits for France, Europe, and Global. Start times are inclusive and end times exclusive; all times are UTC. Window counts use a one-hour stride.
Paradigm / branch
Model
Global
Europe
France
Backbone parameters
SSSV
xPatch
71,076
71,076
71,076
SSSV
PatchTST
455,112
455,112
455,112
SSMV
iTransformer
2,139,976
2,139,976
2,139,976
SSMV
TQNet
596,840
596,840
596,840
MSSV
S 2 Transformer
2,404,936
2,404,936
2,404,936
Appendix
Table 8: Parameter counts across spatial scales. The upper block lists backbone parameters; the lower block lists only additional adapter parameters for the recorded branch configurations.
Paradigm
Model
France
Europe
Global
SSSV
DLinear
<0.001
<0.001
0.0044
xPatch
<0.001
<0.001
0.0058
PatchTST
<0.001
<0.001
<0.001
SSMV
iTransformer
<0.001
<0.001
0.0058
TQNet
<0.001
0.0069
0.0313
DUET
0.0043
0.0193
0.0128
Appendix
Table 9: Seed-level significance tests for the MSE reductions in Table 3 . Entries are two-sided paired t -test p -values after Holm correction within each dataset; bold values are significant at α=0.05 . Each test uses three paired seeds ( df=2 ).
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.3146±0.0003
0.2928±0.0000
−6.94
0.4411±0.0005
0.4259±0.0000
−3.44
xPatch
0.1977±0.0004
0.1921±0.0003
−2.81
0.3371±0.0003
0.3342±0.0003
−0.89
PatchTST
0.1987±0.0005
0.1861±0.0005
−6.36
0.3393±0.0005
0.3298±0.0005
−2.79
Timer
0.3078±0.0000
0.2718±0.0000
−11.67
0.4235±0.0000
0.4034±0.0000
−4.74
TimeMoE
0.2126±0.0000
0.1976±0.0000
−7.07
0.3476±0.0000
0.3377±0.0000
−2.83
Appendix
Table 10: France — Air temperature. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.8359±0.0003
0.7586±0.0000
−9.25
0.6746±0.0014
0.6506±0.0000
−3.56
xPatch
0.7956±0.0006
0.7054±0.0003
−11.34
0.6394±0.0003
0.6171±0.0003
−3.49
PatchTST
0.7811±0.0002
0.7029±0.0001
−10.01
0.6374±0.0004
0.6165±0.0001
−3.28
Timer
0.9058±0.0000
0.7490±0.0000
−17.31
0.6968±0.0000
0.6434±0.0000
−7.68
TimeMoE
0.8696±0.0000
0.7362±0.0000
−15.34
0.6694±0.0000
0.6328±0.0000
−5.47
Appendix
Table 11: France — Wind speed. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.7059±0.0007
0.6835±0.0000
−3.16
0.6742±0.0020
0.6734±0.0000
−0.12
xPatch
0.4959±0.0007
0.4671±0.0006
−5.82
0.5305±0.0005
0.5316±0.0005
+0.22
PatchTST
0.4882±0.0005
0.4657±0.0005
−4.61
0.5346±0.0001
0.5311±0.0003
−0.67
Timer
0.7680±0.0000
0.6644±0.0000
−13.50
0.6638±0.0000
0.6551±0.0000
−1.31
TimeMoE
0.5592±0.0000
0.5141±0.0000
−8.07
0.5644±0.0000
0.5635±0.0000
−0.16
Appendix
Table 12: France — Relative humidity. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.1346±0.0004
0.1069±0.0000
−20.53
0.2715±0.0013
0.2379±0.0000
−12.38
xPatch
0.1005±0.0004
0.0922±0.0003
−8.25
0.2211±0.0006
0.2141±0.0006
−3.16
PatchTST
0.0941±0.0006
0.0856±0.0005
−8.99
0.2121±0.0009
0.2044±0.0008
−3.64
Timer
0.1030±0.0000
0.0929±0.0000
−9.73
0.2222±0.0000
0.2133±0.0000
−4.03
TimeMoE
0.1002±0.0000
0.0901±0.0000
−10.06
0.2174±0.0000
0.2084±0.0000
−4.13
Appendix
Table 13: France — Station pressure. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.2270±0.0003
0.2107±0.0000
−7.15
0.3684±0.0004
0.3536±0.0000
−4.03
xPatch
0.1481±0.0002
0.1445±0.0001
−2.41
0.2861±0.0002
0.2836±0.0002
−0.89
PatchTST
0.1505±0.0006
0.1432±0.0006
−4.84
0.2898±0.0008
0.2836±0.0008
−2.16
Timer
0.2039±0.0000
0.1859±0.0000
−8.83
0.3386±0.0000
0.3263±0.0000
−3.63
TimeMoE
0.1561±0.0000
0.1482±0.0000
−5.09
0.2917±0.0000
0.2858±0.0000
−2.03
Appendix
Table 14: Europe — Air temperature. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.7531±0.0004
0.6782±0.0000
−9.94
0.6265±0.0014
0.6085±0.0000
−2.87
xPatch
0.6890±0.0004
0.6318±0.0002
−8.29
0.5836±0.0002
0.5780±0.0002
−0.96
PatchTST
0.6775±0.0008
0.6272±0.0001
−7.42
0.5830±0.0002
0.5762±0.0001
−1.18
Timer
0.7625±0.0000
0.6617±0.0000
−13.23
0.6238±0.0000
0.5979±0.0000
−4.16
TimeMoE
0.7422±0.0000
0.6554±0.0000
−11.68
0.6048±0.0000
0.5905±0.0000
−2.35
Appendix
Table 15: Europe — Wind speed. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.7719±0.0006
0.7191±0.0000
−6.85
0.6820±0.0019
0.6794±0.0000
−0.39
xPatch
0.5427±0.0001
0.5119±0.0002
−5.68
0.5439±0.0003
0.5416±0.0003
−0.43
PatchTST
0.5371±0.0015
0.5131±0.0016
−4.46
0.5502±0.0015
0.5499±0.0011
−0.05
Timer
0.7665±0.0000
0.6671±0.0000
−12.98
0.6468±0.0000
0.6440±0.0000
−0.43
TimeMoE
0.5974±0.0000
0.5476±0.0000
−8.33
0.5686±0.0000
0.5617±0.0000
−1.22
Appendix
Table 16: Europe — Relative humidity. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.0404±0.0004
0.0268±0.0000
−33.59
0.1505±0.0017
0.1128±0.0000
−25.07
xPatch
0.0238±0.0001
0.0236±0.0001
−0.64
0.1023±0.0003
0.1022±0.0003
−0.06
PatchTST
0.0236±0.0001
0.0229±0.0001
−2.99
0.1017±0.0003
0.1003±0.0003
−1.34
Timer
0.0240±0.0000
0.0232±0.0000
−3.04
0.1016±0.0000
0.1001±0.0000
−1.48
TimeMoE
0.0234±0.0000
0.0227±0.0000
−3.23
0.0986±0.0000
0.0972±0.0000
−1.47
Appendix
Table 17: Europe — Station pressure. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.1837±0.0008
0.1685±0.0011
−8.25
0.3285±0.0009
0.3124±0.0012
−4.88
xPatch
0.1186±0.0010
0.1160±0.0011
−2.25
0.2507±0.0015
0.2495±0.0015
−0.46
PatchTST
0.1188±0.0000
0.1139±0.0001
−4.11
0.2514±0.0001
0.2478±0.0001
−1.45
Timer
0.1628±0.0000
0.1500±0.0000
−7.88
0.2979±0.0000
0.2895±0.0000
−2.80
TimeMoE
0.1258±0.0000
0.1208±0.0000
−3.95
0.2562±0.0000
0.2533±0.0000
−1.12
Appendix
Table 18: Global — Air temperature. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.8466±0.0053
0.7770±0.0010
−8.22
0.6877±0.0024
0.6675±0.0006
−2.94
xPatch
0.7833±0.0041
0.7216±0.0003
−7.87
0.6473±0.0013
0.6348±0.0004
−1.94
PatchTST
0.7767±0.0005
0.7179±0.0002
−7.57
0.6473±0.0002
0.6333±0.0001
−2.17
Timer
0.8832±0.0000
0.7617±0.0000
−13.76
0.6968±0.0000
0.6576±0.0000
−5.62
TimeMoE
0.8533±0.0000
0.7477±0.0000
−12.38
0.6747±0.0000
0.6476±0.0000
−4.01
Appendix
Table 19: Global — Wind speed. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.7290±0.0071
0.6969±0.0040
−4.41
0.6677±0.0031
0.6673±0.0023
−0.05
xPatch
0.5100±0.0042
0.4878±0.0035
−4.36
0.5264±0.0023
0.5214±0.0024
−0.95
PatchTST
0.4980±0.0006
0.4829±0.0008
−3.03
0.5285±0.0009
0.5261±0.0006
−0.46
Timer
0.7494±0.0000
0.6601±0.0000
−11.92
0.6430±0.0000
0.6371±0.0000
−0.91
TimeMoE
0.5753±0.0000
0.5332±0.0000
−7.32
0.5577±0.0000
0.5611±0.0000
+0.61
Appendix
Table 20: Global — Relative humidity. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
MSE ↓
MAE ↓
Model
Base
+Adapter
Δ (%)
Base
+Adapter
Δ (%)
DLinear
0.0438±0.0013
0.0263±0.0000
−39.83
0.1603±0.0019
0.1163±0.0001
−27.50
xPatch
0.0230±0.0004
0.0229±0.0004
−0.59
0.1043±0.0012
0.1043±0.0012
−0.08
PatchTST
0.0224±0.0001
0.0218±0.0001
−2.71
0.1026±0.0004
0.1015±0.0004
−1.05
Timer
0.0232±0.0000
0.0225±0.0000
−2.85
0.1046±0.0000
0.1033±0.0000
−1.22
TimeMoE
0.0227±0.0000
0.0220±0.0000
−3.06
0.1021±0.0000
0.1009±0.0000
−1.16
Appendix
Table 21: Global — Station pressure. MSE and MAE (mean ± sample standard deviation). Δ (%) is signed relative change; negative is better.
Figure 6: Individual-model counterpart of Figure 5 . Open circles denote native predictions, filled circles denote adapted predictions, and squares mark the unadapted HiSTGNN. Thin rules separate SSSV, SSMV, MSSV, and MSMV; lower MSE is better.
France
Model
Base
Interaction only
Full
DLinear
0.4977±0.0004
0.4800±0.0004
0.4605±0.0000
xPatch
0.3974±0.0004
0.3858±0.0003
0.3642±0.0002
PatchTST
0.3905±0.0004
0.3724±0.0004
0.3601±0.0003
Timer
0.5211±0.0000
0.4704±0.0000
0.4445±0.0000
TimeMoE
0.4354±0.0000
0.4028±0.0000
0.3845±0.0000
Appendix
Table 22: Calibration ablation on France and Global. MSE is reported as mean ± sample standard deviation across seeds 2024–2026. Lower is better. Full includes calibration and the same interaction branches. HiSTGNN is excluded.
Backbone
Residual source
MSE ↓
Δ (%)
PatchTST
Base
0.4030±0.0001
—
In-sample full
0.3703±0.0001
−8.11
In-sample matched
0.3790±0.0001
−5.94
Held-out
0.3716±0.0001
−7.79
iTransformer
Base
0.3813±0.0005
—
In-sample full
0.3525±0.0003
−7.54
Appendix
Table 23: Sensitivity to residual fitting sources on France. MSE is reported as mean ± sample standard deviation over three seeds. Δ is the mean per-seed percentage change relative to the shared Base; negative values indicate improvement. Bold marks the lowest mean MSE per backbone.
France
Europe
Paradigm
Backbone
Time (min)
Memory (GiB)
Time (min)
Memory (GiB)
SSSV
xPatch
11.68
0.05
3.02
0.23
PatchTST
9.95
1.37
1.96
7.56
SSMV
iTransformer
6.48
0.77
2.83
4.18
DUET
13.81
0.71
3.15
3.56
MSSV
Corrformer
40.67
26.31
30.93
30.03
Appendix
Table 24: Backbone training cost at batch size 2. Time is minutes per training epoch; memory is peak GPU allocation in GiB.
France
Europe
Frozen backbone
Fit time (min)
Memory (GiB)
Time ratio (%)
Fit time (min)
Memory (GiB)
Time ratio (%)
xPatch
2.77
0.023
23.7
0.64
0.093
21.1
PatchTST
2.33
0.880
23.5
0.80
4.550
40.6
iTransformer
2.93
0.700
45.3
0.62
2.560
22.0
DUET
3.32
0.690
24.1
0.71
2.510
22.6
Corrformer
16.01
28.720
39.4
12.16
21.350
39.3
Appendix
Table 25: Complete fitting cost of the full dual-branch adapter with a frozen backbone. Tfit/Tepoch compares total adapter fitting time with one training epoch of the same backbone.
Figure 7: MSE versus look-back window on France with a fixed 72-hour forecast horizon. Each curve reports the mean MSE across repeated runs for a native backbone without an adapter. HiSTGNN maintains the lowest MSE across all four windows.
Figure 8: Sensitivity of the station-interaction adapter to the KNN neighborhood size K . Each curve shows the three-run mean MSE averaged over four variables; lower is better.
Figure 9: Air temperature forecasts on Europe. Station ITMU0016052; forecast start 2020-12-31 00:00 UTC. Each panel compares a native backbone with its full adapter over the same 72-hour window following 48 hours of history. Adapted forecasts better follow the sustained warming trend.
Figure 10: Wind speed forecasts on Europe. Station NOI0000ENFG; forecast start 2021-12-06 18:00 UTC. Each panel compares a native backbone with its full adapter over the same 72-hour window following 48 hours of history. Adapted forecasts capture the increase in wind speed, while rapid fluctuations remain difficult to reproduce.
Figure 11: Relative humidity forecasts on Europe. Station ROM00015317; forecast start 2021-01-25 06:00 UTC. Each panel compares a native backbone with its full adapter over the same 72-hour window following 48 hours of history. Adapted forecasts better track the humidity decline, with residual differences in timing and magnitude.
Figure 12: Station pressure forecasts on Europe. Station AUM00011126; forecast start 2021-05-27 19:00 UTC. Each panel compares a native backbone with its full adapter over the same 72-hour window following 48 hours of history. Most adapted forecasts follow the observed pressure level and gradual variation more closely; STELLA retains pronounced fluctuations.
Multi-station multivariate weather forecasting aims to forecast future weather variables at multiple weather stations from historical surface observations. Existing station forecasting models learn statistical dependencies among discrete stations, but lack explicit physical evolution. Meanwhile, PDE-based weather models provide interpretable physical dynamics, yet require continuous fields and upper-air variables unavailable in surface station data. To bridge this gap, we propose StationPDE, a station-oriented surface PDE learning model. StationPDE constructs a terrain-aware continuous surface field from discrete station observations and decomposes its physical evolution into surface wind transport and upper-air inference. Surface wind transport explicitly evolves observable weather variables, while upper-air inference uses learnable horizontal diffusion to approximate the missing influence of unavailable upper-air variables. A parallel data-driven diffusion branch captures complementary motion patterns, and an adaptive router integrates the two forecasts for station-level multivariate forecasting. Experiments on Weather2K and MeteoNet show that StationPDE consistently outperforms state-of-the-art baselines, reducing MSE by about 9.6% on average compared with the strongest baseline.
Station-based weather forecasting supports daily life and economic activity, yet accurate forecasts require modeling complex spatial dependencies among stations. Recent clustering-based selective modeling offers a promising alternative to dense inter-station interactions. However, a grouping shared across an observation window may obscure local changes in station relationships, while intra-cluster interactions alone may miss important global context. The theoretical advantages of selective interactions over dense connectivity also remain insufficiently understood. We therefore propose STCFormer, an adaptive spatio-temporal Transformer that dynamically groups stations according to their local evolution within each temporal patch. Its Cluster-Guided Attention Block combines fine-grained local attention within clusters and global attention over regional state summaries, allowing each station to access information beyond its own cluster. We further show that a derived Lipschitz upper bound for cluster-conditioned local attention is no larger than its fully connected counterpart, explaining a potential robustness benefit and motivating the design of InfoLoss. Experiments on three real-world weather datasets spanning eight temperature and wind forecasting tasks show that STCFormer achieves the lowest 24-hour mean squared error on all eight tasks and ranks first or second in 47 of 48 comparisons across metrics and forecasting horizons. Ablations and case studies further confirm the benefits of locally adaptive grouping and complementary local-global interactions. Our code can be obtained at https://github.com/hnu-vis/STCFormer.
Rongwen Li, Haixin Xie, Mingyang Wang +5
Hunan University · China Meteorological Administration · Hunan Provincial Meteorological Bureau
Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices actually improve real-world data assimilation. We present the first controlled benchmark of generative weather data assimilation on real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four weather variables, we evaluate methods while holding the dataset, observation operator, and deep learning architecture fixed. The benchmark compares the major design choices, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three clear conclusions. First, learned generative priors outperform the Gaussian prior of 3D-Var (35.7% vs. 33.3% RMSE reduction over ERA5) despite using no ERA5 background field at inference. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. We further evaluate both dense and sparse station settings and find advantages from generative AI and full-gradient guidance more pronounced under sparsity. Together, these results identify which components of generative weather data assimilation improve performance on real station observations and establish a standardized benchmark for future work.
Ruizhe Huang, Qidong Yang, Jonathan Giezendanner +1
Massachusetts Institute of Technology, Cambridge, MA, USA