Varda-single-1.0: deterministic data-driven weather forecasting at 1 km resolution over Switzerland's complex topography
Authors: Alberto Pennino, Francesco Zanetta, Michele Cattaneo, Claire Merker, Radi Radev, Jonas Bhend, Louis Frey, Hugues de Laroussilhe, +23 more
Organizations: Federal Office of Meteorology and Climatology MeteoSwiss · Swiss Data Science Center (SDSC), ETH Zürich · Center for Climate Systems Modeling (C2SM), ETH Zürich · European Centre for Medium-Range Weather Forecasts (ECMWF) · Norwegian Meteorological Institute (MET Norway) · Royal Netherlands Meteorological Institute (KNMI) · Institute for Atmospheric and Climate Science (IAC), ETH Zürich
We present Varda-single-1.0, a medium-range data-driven weather prediction system built for the Alpine domain. It provides hourly deterministic regional forecasts on a mesh of 1 km resolution and global forecasts on a 31 km mesh. The system comprises two independently trained stretched-grid Graph Transformer models with encoder-processor-decoder architecture, developed in the Anemoi framework: a 6-hourly autoregressive forecaster and a temporal downscaler reconstructing hourly forecasts between the forecaster's steps. Its training curriculum includes pre-training on ERA5 reanalysis data, followed by training on a 20-year kilometre-scale regional reanalysis, and finally fine-tuning on operational kilometre-scale analyses. Verified over one year against operational analyses and surface station observations, Varda-single is competitive with or improves on MeteoSwiss' operational numerical weather prediction baselines for most headline scores and variables. It broadly matches the skill of the high-resolution 1 km ICON-CH1-EPS control at lead times up to +33 h and generally outperforms the 2 km ICON-CH2-EPS control at lead times up to +120 h. Despite competitive aggregate scores, Varda-single underestimates some local wind maxima and produces overly smooth convective precipitation fields, consistent with the smoothing associated with squared-error training. To gain insight into the model's behaviour, we investigate three case studies beyond the aggregated headline scores, and find particular weaknesses in Varda-single's representation of local winds over complex terrain. Varda-single represents an important step in the development of high-resolution ML forecasting over complex terrain, in complementing the operational regional numerical weather prediction models of MeteoSwiss with data-driven models and in providing a pretrained model for researchers and user-specific applications.
Figures & tables
Figure 1: Example of a Varda-single forecast initialised at 00 UTC on 6 April 2025, showing 2 m temperature and 10 m wind speed at a lead time of +12 h .
Description
Type
Model
Single-level features
2t
2 m air temperature
Prog.
F,D
2d
2 m dew point temperature
Prog.
F,D
10u
10 m eastward wind
Prog.
F,D
10v
10 m northward wind
Prog.
F,D
lsm
Land-sea mask
Forc. ⋆
F,D
Table 1: Summary of features included in Varda-single. Variables used by the 6-hourly forecaster are marked with "F", while variables used by the temporal downscaler are marked with "D".
Figure 2: The Varda-single system. The forecaster model takes in the current and previous atmospheric states xt,t−6 and propagates them forward in time by 6 hours to produce xt+6 . The temporal downscaler model takes in inputs xt and xt+6 and produces hourly outputs from xt+1 to xt+5 , also including xt+6 for hourly accumulated variables, in our case total precipitation.
Figure 3: Input meshes of Varda-single. Blue nodes are obtained from the native ERA5 N320 reduced Gaussian grid and represent the spatial boundary conditions necessary to model the regional domain. Red nodes are the cell centres of the REA-L-CH1 grid at 1 km spatial resolution, our regional domain of interest. During pre-training, only global nodes are considered to learn global weather dynamics. In the stretched-grid training step, the regional domain is embedded within the global grid, and regional dynamics are learned conditional on the global dynamics. We display a close-up of a 24 km× 24 km cutout to depict the complex topography of the Alpine region, which spans a substantial portion of the regional domain.
Training
Global
Regional
Training
Steps
GPUs
Global
LR
Rollout
stage
dataset
dataset
period
batch size
length
6-hour forecaster
Global pre-training
ERA5 N320
–
1979–2023 (45 yr)
50,000
32
32
10−5
1
Stretched-grid training
ERA5 N320
REA-L-CH1
2005–2023 (19 yr)
100,000
32
16
6×10−6
1
Rollout training
ERA5 N320
REA-L-CH1
2005–2023 (19 yr)
8,000
64
16
6×10−6
2–6
Operational fine-tuning
IFS N320
KENDA-CH1
2024–2025 (2 yr)
1,000
64
16
10−6
6
Table 2: Curriculum training stages for the 6-hourly forecaster and the hourly temporal downscaler. LR denotes the peak learning rate. For the forecaster, a cosine schedule with 10 warm-up steps is used. For the temporal downscaler, the learning rate is linearly warmed up over 100 steps and subsequently annealed with a cosine schedule to 3×10−7 at the final step. The effective number of samples seen by the model is the number of steps multiplied by the global batch size.
Figure 4: Forecast performance by forecast lead time for Varda-single and the operational baseline forecasts from ICON and AIFS v1.1. RMSE (top row), bias (second row) and ETS (bottom rows) have been computed against the operational KENDA-CH1 analysis and averaged across all initialisations in the test period and all grid points in the inner verification domain. Forecasts from the 6-hourly forecaster are indicated with dots.
Figure 5: An example of a single 6-hourly forecasting step of Varda-single initialised at 06 UTC on 6 April 2026. The left column shows the model output for 2 m temperature, 10 m wind and 2 m dew point temperature; the middle column shows the corresponding operational KENDA-CH1 analysis. The right column shows the pixel-wise difference between the forecast and the analysis.
Figure 6: Station-based scorecards: relative percentage difference for RMSE and ETS comparing Varda-single with ICON-CH1-CTRL (panel a) and ICON-CH2-CTRL (panel b). Blue dots: Varda-single better than the baseline NWP. Red dots: NWP baseline better than Varda-single. The size of the dots scales with the magnitude of the relative difference.
Figure 7: Pointwise MSE skill score of Varda-single against the ICON-CH1-CTRL baseline at +6 h lead time, accumulated over all initialisations and verified against the KENDA-CH1 analysis. Top row: 2 m temperature; bottom row: 10 m wind speed. Left column: total skill score S ( Equation 1 ). Right column: the contribution SBIAS of the systematic error; the random contribution SSTDE is the difference between the two columns. Blue indicates that Varda-single outperforms the baseline, red that it underperforms. Both columns share the colour scale, which saturates beyond ±0.65 .
Figure 8: Meteograms for the forecast initialised at 18 UTC on 27 June 2025 for 2 m temperature and 10 m wind speed and direction. The left column shows the valley station in Sion (SIO) and the right column the flat-region reference in Kloten (KLO), for SwissMetNet station observations (green), ICON-CH1-EPS (light blue) and ICON-CH2-EPS (dark blue) ensemble means, and Varda-single (red).
Figure 9: Meteograms for the forecast initialised at 00 UTC on 21 March 2025 for 2 m temperature (a), 10 m wind direction (b) and speed (d) in Altdorf (ALT), and the mean sea-level pressure difference between ALT and Lugano (LUG, c), for SwissMetNet station observations (green), ICON-CH1-CTRL (light blue) and ICON-CH2-CTRL (dark blue), and Varda-single (red).
Figure 10: 6 h accumulated precipitation for a well-forecast winter case (top; initialised at 00 UTC on 7 December 2025, lead time +24 h ) and a poorly forecast summer case (bottom; initialised at 12 UTC on 2 July 2025, lead time +6 h ). Left: Varda-single. Right: the verifying KENDA-CH1 analysis. Inset boxes give the SAL components for the displayed 6 h window. Note that these are single-window values, whereas the values in Figure 11 are per-initialisation averages over the 6–30 h windows and are therefore numerically different.
Figure 11: SAL statistics over a subset of the test period for (a) Varda-single and (b) ICON-CH1-CTRL. Each circle is one initialisation, positioned by its mean structure ( S ) and amplitude ( A ) components over the 6–30 h windows, with the circle area scaling with the mean location component ( L ); diamonds mark the seasonal means. Only initialisations with at least 2 mm of domain-mean precipitation in the KENDA-CH1 analysis over the scored window are shown (34 in JJA 2025, 35 in DJF 2025/26). The two events shown in Figure 10 are labelled.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Graph connectivity shown across three representative regions (left to right in each row: global ERA5 region, boundary between global and LAM, LAM-only region). Blue dots are data nodes, red dots are hidden nodes, and grey lines are edges.
Figure A2: Map of Switzerland showing the Jura Mountains, the Swiss Plateau and the Swiss Alps, together with the surface stations used in the evaluation. There are a total of 187 stations measuring 2 m temperature, 151 stations measuring 10 m wind speed and direction, 59 stations measuring mean sea-level pressure and 277 stations measuring total precipitation.
Accurate and timely weather forecasts are critical for high-impact decisions in modern society. Machine-learning-based weather prediction is emerging as an alternative for producing initial conditions, forecasts, and even both in end-to-end systems. These methods deliver predictions faster and often with higher skill than traditional numerical weather prediction (NWP). However, even end-to-end models typically rely on NWP-generated reanalyses for supervision, thereby inheriting the biases and resolution limitations of those NWPs, and limiting adaptation to settings where suitable reanalysis products are unavailable, infrequently updated, or expensive to produce. Here we introduce ObsCast, a regional system that generates both analysis and predictions, without using any NWP-derived data in either training or inference, while still achieving state-of-the-art performance in short-term high-resolution regional modeling. Over the contiguous United States and Europe, ObsCast outperforms operational NWP for near-surface variables through 18 h and produces skillful precipitation forecasts. It provides a simpler and more adaptable route to build and refine regional forecasting services directly from local observations, without the need to develop complex and costly traditional forecasting pipelines.
Pengcheng Zhao, Siqi Xiang, Weixin Jin +10
1Microsoft Corporation. · University of Cambridge, Cambridge, United Kingdom. · 3The Alan Turing Institute, London, United Kingdom.
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best physics-based models and they are exclusively initialized with and trained on analysis data. As a result, they cannot directly make use of observations, and any biases in the analysis are inherited by the forecast. WeatherNext 3 addresses these shortcomings and establishes a new state-of-the-art for probabilistic medium-range forecasting skill. First, WeatherNext 3 generates new forecasts every hour (rather than every 6 hours like traditional global models) by ingesting low-latency geostationary satellite data. Second, WeatherNext 3's temporal and spatial resolution are on par with physics-based global models, with hourly time steps and 0.1 degree resolution for single-level variables, including solar radiation and cloud cover. Third, WeatherNext 3 moves beyond traditional analysis variables by learning to predict satellite-derived precipitation estimates, as well as tropical cyclone and station observations. Modelling sparse station data allows WeatherNext 3 to make 2m temperature and dewpoint predictions at any location and time, conditioned on local geographical features, with substantially lower error than competing global models, even when evaluated against unseen stations. Together, WeatherNext 3's capabilities move operational AI-based weather forecasting beyond emulating the traditionally distinct stages of data assimilation, forecasting and post-processing, which helps to further push the frontier of performance and granularity for global weather prediction.
Multi-station multivariate weather forecasting aims to forecast future weather variables at multiple weather stations from historical surface observations. Existing station forecasting models learn statistical dependencies among discrete stations, but lack explicit physical evolution. Meanwhile, PDE-based weather models provide interpretable physical dynamics, yet require continuous fields and upper-air variables unavailable in surface station data. To bridge this gap, we propose StationPDE, a station-oriented surface PDE learning model. StationPDE constructs a terrain-aware continuous surface field from discrete station observations and decomposes its physical evolution into surface wind transport and upper-air inference. Surface wind transport explicitly evolves observable weather variables, while upper-air inference uses learnable horizontal diffusion to approximate the missing influence of unavailable upper-air variables. A parallel data-driven diffusion branch captures complementary motion patterns, and an adaptive router integrates the two forecasts for station-level multivariate forecasting. Experiments on Weather2K and MeteoNet show that StationPDE consistently outperforms state-of-the-art baselines, reducing MSE by about 9.6% on average compared with the strongest baseline.