Multi-station multivariate weather forecasting aims to forecast future weather variables at multiple weather stations from historical surface observations. Existing station forecasting models learn statistical dependencies among discrete stations, but lack explicit physical evolution. Meanwhile, PDE-based weather models provide interpretable physical dynamics, yet require continuous fields and upper-air variables unavailable in surface station data. To bridge this gap, we propose StationPDE, a station-oriented surface PDE learning model. StationPDE constructs a terrain-aware continuous surface field from discrete station observations and decomposes its physical evolution into surface wind transport and upper-air inference. Surface wind transport explicitly evolves observable weather variables, while upper-air inference uses learnable horizontal diffusion to approximate the missing influence of unavailable upper-air variables. A parallel data-driven diffusion branch captures complementary motion patterns, and an adaptive router integrates the two forecasts for station-level multivariate forecasting. Experiments on Weather2K and MeteoNet show that StationPDE consistently outperforms state-of-the-art baselines, reducing MSE by about 9.6% on average compared with the strongest baseline.
Figures & tables
Figure 1 : Motivation of station-oriented surface PDE modeling. (a) Wind-aligned redistribution of surface relative humidity observed from MeteoNet stations. (b) Surface PDE approximation under unavailable upper-air dynamics, where surface wind transport and upper-air influence are modeled on the station-derived surface field.
Figure 2 : The overview of the StationPDE structure.
Models
StationPDE (Ours)
CDPNet
TimeFilter
MultiPatch Former
TimeXer
TimeMixer
iTransformer
Corrformer
DLinear
weather2k
pressure
4.413 ±0.090
9.840 ±0.644
8.625 ±0.137
6.750 ±0.122
7.738 ±0.080
7.684 ±0.119
7.700 ±0.077
5.732 ±0.047
11.317 ±0.129
temp
4.607 ±0.055
5.856 ±0.101
7.058 ±0.022
6.694 ±0.038
7.121 ±0.020
7.653 ±0.020
7.118 ±0.015
5.897 ±0.016
8.356 ±0.006
humidity
127.240 ±0.807
150.312 ±2.075
171.486 ±0.824
168.889 ±0.440
165.972 ±0.204
174.633 ±0.756
171.770 ±0.373
148.102 ±0.213
187.482 ±0.078
u-wind
2.155 ±0.007
2.234 ±0.011
2.597 ±0.002
2.573 ±0.004
2.516 ±0.005
2.581 ±0.006
2.606 ±0.003
2.374 ±0.003
2.747 ±0.001
v-wind
2.194 ±0.008
2.322 ±0.013
2.820 ±0.010
2.789 ±0.009
2.768 ±0.005
2.864 ±0.013
2.863 ±0.006
2.534 ±0.003
3.083 ±0.001
metonet
temp
4.298 ±0.095
4.756 ±0.077
5.348 ±0.069
5.310 ±0.030
5.389 ±0.067
6.298 ±0.037
5.314 ±0.039
7.444 ±0.164
6.762 ±0.046
Table 1 : Experiment results (MSE) on the Weather2K and MeteoNet datasets. All models use 48 hours of historical observations to forecast the next 24 hours. Best results are marked in bold , and second-best results are underlined . Complete MAE results are reported in the appendix.
Models
StationPDE (Ours)
CDPNet
TimeFilter
MultiPatch Former
TimeXer
TimeMixer
iTransformer
Corrformer
DLinear
temp
48
6.978 ±0.179
7.288 ±0.057
8.556 ±0.183
8.397 ±0.120
8.483 ±0.055
9.274 ±0.091
8.323 ±0.075
10.449 ±0.374
9.969 ±0.136
96
9.849 ±0.278
10.685 ±0.144
12.891 ±0.154
12.845 ±0.148
12.726 ±0.118
13.352 ±0.065
12.761 ±0.051
14.452 ±0.188
13.940 ±0.063
humidity
48
103.910 ±3.200
107.508 ±0.805
122.763 ±1.142
117.479 ±0.315
118.568 ±0.778
126.363 ±0.822
120.362 ±0.260
135.347 ±1.055
122.470 ±0.283
96
133.674 ±2.090
138.263 ±0.455
146.440 ±2.454
141.550 ±0.817
142.138 ±1.698
148.828 ±0.826
144.545 ±0.504
155.502 ±1.214
139.269 ±0.108
u-wind
48
5.794 ±0.081
5.865 ±0.039
6.669 ±0.118
6.686 ±0.020
6.630 ±0.014
7.137 ±0.018
6.724 ±0.025
7.138 ±0.076
6.669 ±0.025
96
7.174 ±0.225
7.219 ±0.037
8.433 ±0.172
8.453 ±0.033
8.513 ±0.032
8.970 ±0.024
8.480 ±0.028
9.028 ±0.307
7.881 ±0.007
Table 2 : Experiment results (MSE) on the MeteoNet dataset with different forecast horizons (48h and 96h). All models use 48 hours of historical observations. Best results are marked in bold , and second-best results are underlined . Complete MAE results are reported in the appendix.
Method
pressure
temp
humidity
u-wind
v-wind
Full / StationPDE
4.41 ±0.09
4.61 ±0.06
127.24 ±0.81
2.16 ±0.01
2.19 ±0.01
w/o PDE
4.56 ±0.03
5.42 ±0.06
150.62 ±1.55
2.39 ±0.01
2.41 ±0.00
w/o Data-Driven
4.84 ±0.06
5.87 ±0.11
161.17 ±0.70
2.41 ±0.06
2.42 ±0.01
w/o Station Residual
4.72 ±0.06
5.36 ±0.11
148.95 ±0.74
2.38 ±0.01
2.40 ±0.00
w/o Lobs
4.45 ±0.07
4.76 ±0.01
134.88 ±0.61
2.16 ±0.01
2.20 ±0.01
Table 3 : Ablation study results (MSE) on the Weather2K dataset. All models use 48 hours of historical observations to forecast the next 24 hours. Best results are marked in bold based on higher-precision values.
Station weather forecasting is fundamentally shaped by both complex spatial dependencies across stations and strong physical coupling among weather variables. However, existing studies often consider these relationships separately and use different datasets and experimental settings, hindering systematic assessment of their individual and joint contributions. In this paper, we introduce M2Weather, a benchmark for joint multi-station and multi-variable weather forecasting. Through multi-criteria quality control and station stratification, we collect 2,809 high-quality stations with 5 physically coupled weather variables across three spatial scales: France, Europe, and Global. This multi-scale design lets us examine whether conclusions persist from national to global station networks. We also introduce unified training and evaluation protocols to enable fair comparison of different station-variable modeling paradigms. To further examine the benefits of modeling station-variable relationships, we design a lightweight, plug-and-play adapter. With a trained weather forecasting model, this adapter can introduce missing station or variable relationships without retraining the model. This enables fair and efficient investigation of station-variable relationships. Systematic evaluation of 16 representative models shows the benefits of jointly modeling station and variable relationships. Completing missing relationships further reduces MSE for all adapted models on all three datasets. Together, these results identify the complementary information across stations and variables as an important resource for improving station weather forecasting. Our code can be obtained at https://github.com/hnu-vis/M2-Weather.
Rongwen Li, Xiao Wang, Mingyang Wang +4
Hunan University · China Meteorological Administration
Near-surface weather varies over tens to hundreds of meters, yet remains unresolved in analyses and forecasts. We test whether this variation can be inferred without resolving atmospheric dynamics. Combining sparse weather stations, high-resolution Earth observation, and coarse atmospheric dynamics, we infer temperature, dewpoint, and wind at 30-m resolution across the contiguous United States. Against measurements held out in space and time, estimates reduce error by 11-28% relative to the strongest baseline. Within held-out 0.25∘ grid cells, we recover more spatial variance than baselines, explaining nearly half of temperature variability in the median cell. The method captures time-varying differences between locations and produces coherent patterns associated with topography and land cover. Beyond weather, our findings illustrate how sparse observations of a dynamical system can be combined with dense observations of persistent environmental structure to recover otherwise unresolved spatial variability.
Jonathan Giezendanner, Qidong Yang, Ruizhe Huang +7
Massachusetts Institute of Technology · Shell Information Technology · Allen Institute for AI +1
We present Varda-single-1.0, a medium-range data-driven weather prediction system built for the Alpine domain. It provides hourly deterministic regional forecasts on a mesh of 1 km resolution and global forecasts on a 31 km mesh. The system comprises two independently trained stretched-grid Graph Transformer models with encoder-processor-decoder architecture, developed in the Anemoi framework: a 6-hourly autoregressive forecaster and a temporal downscaler reconstructing hourly forecasts between the forecaster's steps. Its training curriculum includes pre-training on ERA5 reanalysis data, followed by training on a 20-year kilometre-scale regional reanalysis, and finally fine-tuning on operational kilometre-scale analyses. Verified over one year against operational analyses and surface station observations, Varda-single is competitive with or improves on MeteoSwiss' operational numerical weather prediction baselines for most headline scores and variables. It broadly matches the skill of the high-resolution 1 km ICON-CH1-EPS control at lead times up to +33 h and generally outperforms the 2 km ICON-CH2-EPS control at lead times up to +120 h. Despite competitive aggregate scores, Varda-single underestimates some local wind maxima and produces overly smooth convective precipitation fields, consistent with the smoothing associated with squared-error training. To gain insight into the model's behaviour, we investigate three case studies beyond the aggregated headline scores, and find particular weaknesses in Varda-single's representation of local winds over complex terrain. Varda-single represents an important step in the development of high-resolution ML forecasting over complex terrain, in complementing the operational regional numerical weather prediction models of MeteoSwiss with data-driven models and in providing a pretrained model for researchers and user-specific applications.
Alberto Pennino, Francesco Zanetta, Michele Cattaneo +28
Federal Office of Meteorology and Climatology MeteoSwiss · Swiss Data Science Center (SDSC), ETH Zürich · Center for Climate Systems Modeling (C2SM), ETH Zürich +4