Recent probabilistic weather forecasters train stochastic predictors with the continuous ranked probability score (CRPS) to generate each ensemble member in a single forward pass. These models learn the predictive distribution from the forecast context alone, which becomes difficult at longer forecast horizons where uncertainty is high. To learn the predictive distribution more effectively, we introduce auxiliary conditional denoising tasks that predict the same future state from the context and its corrupted version, which provides partial future information that can reduce prediction ambiguity. Building on distributional diffusion models, we learn the conditional distributions of these tasks with a single stochastic predictor by minimizing a proper scoring rule across noise levels. At inference, the predictor can still generate each ensemble member in a single forward pass at the fully corrupted endpoint. Standard CRPS training is recovered as the endpoint-only special case of our formulation, so our framework extends existing CRPS-based forecasters with only additional conditioning inputs. Controlled experiments show that the auxiliary tasks improve one-step forecasting across architectures, with larger gains at longer forecast horizons. The gains extend to high-dimensional global weather forecasting under both training from scratch and fine-tuning, along with improved calibration and potential benefits for generalization under distribution shift.
Figures & tables
Figure 1: Comparison of probabilistic weather forecasting methods . Diffusion-based models (left) train a deterministic denoiser, while CRPS-based models (middle) train a stochastic predictor to generate ensembles. Our method (right) retains the stochastic predictor and adds a sampled corrupted future and its noise level as conditioning inputs ( purple ). These inputs introduce auxiliary conditional denoising tasks learned through shared parameters, while preserving one-step ensemble generation.
Figure 2: FCN3 ( Bonev et al., 2025 ) train and validation normalized CRPS (nCRPS) across forecasting horizons.
Figure 3: Schematic comparison of the context-only forecasting task and auxiliary conditional denoising tasks.
Figure 4: Effect of conditional denoising tasks on one-step forecasting . For 3-day-ahead weather forecasting, corrupted futures with noise levels t∈[tmin,1] serve as additional training inputs. Panels show validation nCRPS (left), spread–skill ratio (SSR, middle), and test nCRPS change (%) relative to the CRPS-trained baseline (right). Endpoint-only training ( tmin=1 ) closely matches the CRPS baseline. Lowering tmin broadens the range of auxiliary tasks, mitigating performance degradation.
Figure 5Figure 6
CRPS ↓
RMSE ↓
1d
3d
10d
1d
3d
10d
Method
Z500
T850
Q700
Z500
T850
Q700
Z500
T850
Q700
Z500
T850
Q700
Z500
T850
Q700
Z500
T850
Q700
IFS-ENS †
24.1
0.370
0.283
59.3
0.550
0.427
263.8
1.343
0.725
45.8
0.718
0.611
133.4
1.105
0.902
623.1
2.819
1.440
GenCast †
20.2
0.269
0.222
54.3
0.459
0.367
253.6
1.281
0.676
39.3
0.542
0.498
123.5
0.955
0.807
606.5
2.737
1.375
NeuralGCM †
22.9
0.333
0.263
54.9
0.497
0.380
253.4
1.293
0.679
44.0
0.658
0.543
126.2
1.027
0.812
606.8
2.756
1.374
MOSAIC
22.7
0.300
0.239
58.3
0.480
0.372
261.0
1.303
0.680
44.2
0.601
0.525
133.2
1.001
0.814
619.9
2.781
1.383
Table 3: CRPS and ensemble-mean RMSE across representative variables and forecast lead times 3 3 3 Results marked with † are provided for reference, as their configurations differ from our controlled experimental setup. Our primary comparisons evaluate standard CRPS and DDM training under matched settings. .
Figure 7: Relative CRPS difference between DDM- and CRPS-trained MOSAIC across all 84 forecast channels over 10-day rollouts. Blue indicates better performance of DDM.
Figure 8: Ensemble calibration analysis. (Top left) SSR versus forecast lead time for Z500, T850, and Q700. (Bottom left) SSR distributions across all 84 forecast channels at 1-, 3-, and 10-day lead times. (Right) Binned spread–skill plots for the same variables at 1-, 3-, and 10-day lead times.
nCRPS ↓
nRMSE ↓
SSR dev. ↓
Objective
Relative (%)
Relative (%)
Absolute
CRPS
+3.10
+2.68
+0.017
DDM
+2.60
+2.24
+0.007
Table 4: Changes in performance under the ERA5-to-HRES distribution shift.
Figure 9: (Left) Relative CRPS differences from IFS-ENS for ERA5-trained MOSAIC evaluated on ERA5 and HRES-fc0. (Right) Ensemble track forecasts and IBTrACS best track for Hurricane Ian.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Table 12
# Parameters
Model
Target
Ensemble randomness
CRPS
DDM
Δ
Swin
residual
latent vector ( d=32 )
26.82M
26.84M
+0.09%
U-Net
residual
latent vector ( d=32 )
22.75M
22.76M
+0.05%
FGN
residual
latent vector ( d=32 )
40.71M
40.80M
+0.21%
FCN3
state
internal multi-scale noise
17.84M
17.85M
+0.03%
Appendix
Table 7: Backbones used in the 2.8125∘ pilot study.
Setting
Value
Optimizer
AdamW
Learning rate
1×10−4
Weight decay
1×10−4
Epochs
50
Batch size
16
Gradient clipping
1.0
Appendix
Table 8: Optimization settings, shared by every 2.8125∘ pilot run.
# Parameters
Model
Target
Ensemble randomness
CRPS
DDM
Δ
MOSAIC
state
latent vector ( d=32 )
214.19M
221.11M
+3.23%
U-Cast
residual
Monte Carlo dropout
895.43M
904.96M
+1.06%
Appendix
Table 9: Backbones used in the 1.5∘ experiments.
MOSAIC
U-Cast
Training budget
250,000 steps
100 pretraining / 8 fine-tuning epochs
Effective batch size
16
48
Optimizer
Muon
Muon / AdamW
Peak learning rate
10−3
7×10−3/7×10−5
Learning-rate schedule
cosine
linear warmup + cosine
Final learning rate
10−6
10−8
Appendix
Table 10: Training configurations at 1.5∘ .
Figure 10: Training loss (solid, right axis) and validation nCRPS (dashed, left axis) of CRPS-trained FGN, FCN3, Swin, and U-Net across forecast horizons at 2.8125∘ . Training loss uses residual normalization for FGN, Swin, and U-Net and state normalization for FCN3, whereas validation nCRPS uses state normalization for all models. Training and validation losses therefore share the same normalization only for FCN3. Training configurations are provided in Section C.4 .
Figure 11: Effect of conditional denoising tasks on FGN one-step forecasting. FGN is trained for 100 epochs for 3-day-ahead forecasting at 2.8125∘ . Panels show validation nCRPS (left), validation SSR (middle), and test nCRPS changes (%) relative to the CRPS baseline (right). Test results use checkpoints selected by validation nCRPS, with NFE=1 for all DDM configurations.
Objective
Training t
nCRPS ↓
nRMSE ↓
CRPS
–
0.1915
0.3920
DDM
t=1
0.1906 (−0.47%)
0.3906 (−0.36%)
t∼U[0.75,1]
0.1853 (−3.21%)
0.3822 (−2.52%)
t∼U[0.50,1]
0.1809 (−5.53%)
0.3749 (−4.37%)
t∼U[0.25,1]
0.1780 (−7.03%)
0.3699 (−5.65%)
t∼U[0,1]
0.1759 (−8.10%)
0.3664 (−6.55%)
Appendix
Table 11: Effect of tmin on FCN3- 3d one-step forecasting performance at 2.8125∘ used in Figure 4 .
Objective
Training t
nCRPS ↓
nRMSE ↓
CRPS
–
0.1868
0.3847
DDM
t=1
0.1861 (−0.41%)
0.3828 (−0.49%)
t∼U[0.75,1]
0.1781 (−4.67%)
0.3685 (−4.20%)
t∼U[0.50,1]
0.1736 (−7.11%)
0.3611 (−6.13%)
t∼U[0.25,1]
0.1687 (−9.73%)
0.3528 (−8.27%)
t∼U[0,1]
0.1682 (−9.96%)
0.3517 (−8.55%)
Appendix
Table 12: Effect of tmin on FGN- 3d -100ep one-step forecasting at 2.8125∘ used in Figure 11 . Note that the best checkpoint of the CRPS baseline is the same as FGN- 3d -50ep due to overfitting.
nCRPS ↓
nRMSE ↓
Model
Objective
12h
1d
3d
5d
12h
1d
3d
5d
Swin
CRPS
0.0491
0.0712
0.1839
0.2498
0.0966
0.1429
0.3789
0.5151
DDM
0.0480
0.0698
0.1628
0.2264
0.0947
0.1406
0.3409
0.4739
Δ (%)
-2.1
-2.0
-11.5
-9.4
-1.9
-1.6
-10.0
-8.0
U-Net
CRPS
0.0547
0.0799
0.1759
0.2333
0.1072
0.1581
0.3608
0.4782
DDM
0.0543
0.0789
0.1687
0.2226
0.1064
0.1579
0.3507
0.4632
Appendix
Table 13: Direct forecasting performance at 2.8125∘ across four architectures.
Model
RI
Objective
24h
72h
120h
240h
288h
360h
FGN
12h
CRPS
0.0747
0.1420
0.1928
0.2442
0.2514
0.2571
DDM
0.0743
0.1419
0.1926
0.2443
0.2514
0.2570
Δ (%)
-0.5
-0.1
-0.1
+0.0
+0.0
-0.0
1d
CRPS
0.0792
0.1434
0.1934
0.2444
0.2509
0.2566
DDM
0.0778
0.1412
0.1910
0.2428
0.2499
0.2559
Δ (%)
-1.8
-1.6
-1.2
-0.7
-0.4
-0.3
Appendix
Table 14: Autoregressive rollout nCRPS at 2.8125∘ .
Model
RI
Objective
24h
72h
120h
240h
288h
360h
FGN
12h
CRPS
0.1497
0.2986
0.4060
0.5082
0.5220
0.5332
DDM
0.1490
0.2985
0.4056
0.5085
0.5223
0.5329
Δ (%)
-0.4
-0.0
-0.1
+0.1
+0.1
-0.0
1d
CRPS
0.1586
0.3012
0.4066
0.5083
0.5210
0.5321
DDM
0.1559
0.2978
0.4031
0.5062
0.5199
0.5314
Δ (%)
-1.7
-1.1
-0.9
-0.4
-0.2
-0.1
Appendix
Table 15: Autoregressive rollout nRMSE at 2.8125∘ . The same runs and conventions as Table 14 .
Figure 12: Effective noise levels by harmonic degree for κ0=0.01 (top) and κ0=0.05 (bottom).
CRPS
DDM
Model
RI
baseline
grid
SpG
HK κ0=0.01
HK κ0=0.05
Δ (%)
Swin
12h
0.0491
0.0489
0.0496
0.0484
0.0480
-2.1
1d
0.0712
0.0712
0.0708
0.0698
0.0704
-2.0
3d
0.1839
0.1644
0.1628
0.1645
0.1705
-11.5
5d
0.2498
0.2264
0.2268
0.2297
0.2358
-9.4
U-Net
12h
0.0547
0.0553
0.0561
0.0548
0.0543
-0.7
Appendix
Table 16: Noise scheduling for each backbone and RI at 2.8125∘ , with the single-step nCRPS. Values are test nCRPS. Shaded cells denote the configuration selected by the lowest validation nCRPS.
Figure 13: Relative performance differences in CRPS (left) and ensemble-mean RMSE (right) between DDM and CRPS training for MOSAIC (top) and U-Cast (bottom). Negative values (blue) indicate that DDM shows better performance than CRPS.
Figure 14: CRPS over 10-day autoregressive rollouts in ERA5 2020 .
Figure 15: Ensemble-mean RMSE over 10-day autoregressive rollouts in ERA5 2020 .
Figure 16: SSR over 10-day autoregressive rollouts in ERA5 2020 .
Figure 17: Effect of NFE on MOSAIC forecasting performance. Relative CRPS differences between DDM and CRPS-trained MOSAIC across 84 forecast channels over 10-day autoregressive rollouts. The same DDM checkpoint is evaluated with NFE=1 (left) and NFE=2 (right) in ERA5 2020 .
Figure 18: CRPS over 10-day autoregressive rollouts in HRES-fc0 2020 .
Figure 19: Ensemble-mean RMSE over 10-day autoregressive rollouts in HRES-fc0 2020 .
Figure 20: SSR over 10-day autoregressive rollouts in HRES-fc0 2020 .
Figure 21: Hurricane Michael track forecasts from 50-member MOSAIC (top) and U-Cast (bottom) ensembles, comparing CRPS (blue, left) and DDM (purple, right). Michael occurred in 2018, within the validation period, and is included as an illustrative case study.
Hurricane Ian
Hurricane Michael
Track Error ↓
Track CRPS ↓
Track Error ↓
Track CRPS ↓
Method
Objective
1d
3d
5d
1d
3d
5d
1d
2d
3d
4d
1d
2d
3d
4d
MOSAIC
CRPS
73
89
189
51
68
125
42
81
156
231
30
54
110
153
MOSAIC
DDM
38
44
55
26
39
63
25
17
68
138
17
17
47
86
U-Cast
CRPS
39
59
245
30
43
148
36
75
150
308
25
51
112
231
U-Cast
DDM
35
37
210
24
38
123
35
62
107
210
24
40
72
141
Appendix
Table 17: Tropical-cyclone ensemble-mean track error and track CRPS (km).
Method
Objective
Lead time
150 km
200 km
300 km
MOSAIC
CRPS
48 h
11/50 (22%)
31/50 (62%)
48/50 (96%)
MOSAIC
DDM
48 h
23/50 (46%)
47/50 (94%)
50/50 (100%)
U-Cast
CRPS
120 h
13/50 (26%)
18/50 (36%)
28/50 (56%)
U-Cast
DDM
120 h
16/50 (32%)
23/50 (46%)
31/50 (62%)
Appendix
Table 18: Number and percentage of ensemble members within 150, 200, and 300 km of the time-matched IBTrACS center for Hurricane Ian. MOSAIC is evaluated at 48 h and U-Cast at 120 h, corresponding to the inset lead times as shown in Figures 22 – 24 .
Figure 22: 50-member ensemble track forecasts for Hurricane Ian initialized at 12 UTC on 23 September 2022, from MOSAIC (top) and U-Cast (bottom), comparing CRPS (blue, left) and DDM (purple, right). Insets show the 48 h MOSAIC and 120 h U-Cast forecast positions. Dashed circles mark a 150 km radius around the time-matched IBTrACS center.
Figure 23: Same as Figure 22 , but with a 200 km radius.
Figure 24: Same as Figure 22 , but with a 300 km radius.
Probabilistic weather forecasting requires not only accurate trajectories, but calibrated distributions over plausible atmospheric futures. Recent data-driven systems have achieved remarkable deterministic skill, and diffusion-based ensemble forecasters have substantially improved sample realism and uncertainty quantification. However, their inference cost scales with forecast horizon, ensemble size, and the number of denoising steps required for each transition, making large operational ensembles expensive. To address this, we present Tyche, a one-step conditional flow model for efficient probabilistic weather forecasting. Tyche models the conditional forecast distribution with a destination-aware average-velocity flow that maps Gaussian noise directly to future weather states in a single function evaluation (1-NFE). To make this one-step transport learnable in high-dimensional geophysical fields, we derive a JVP-regularized rectification objective that enforces temporal self-consistency across source and destination flow timesteps without explicitly forming Jacobians. The transport field is parameterized by an isotropic Swin-style transformer that preserves fine-scale spatial structure while remaining scalable on global grids. To improve ensemble reliability under autoregressive forecasting, we further introduce a rollout-based finetuning stage with curriculum CRPS calibration supervision. Experiments on ERA5 at 1.5∘ and 6-hour resolution show that our Tyche, using merely a single NFE, matches or exceeds the forecast skill and calibration of state-of-the-art multi-step generative baselines and the operational ECMWF IFS ensemble.
Deep learning has revolutionised weather forecasting in recent years, especially through atmospheric foundation models, which offer competitive skill for a fraction of the computational costs of classic physics-based models. However, most existing foundation models are deterministic, limiting the generation of large ensembles for accurate uncertainty quantification, extreme weather risk assessment, and long-range weather forecasting. Furthermore, these models incur a large, often prohibitive, computational overhead to train from scratch. To address these shortcomings, we turn a pretrained deterministic prior model, namely the Aurora foundation model, into a generative ensemble-prediction model. To that end, we introduce a novel generative method, Denoising Stochastic Interpolants, combined with a replay buffer for Stochastic Differential Equation (SDE) rollout, enabling probabilistic training of SDE trajectories. Our stochastic foundation model, Xaurora, is finetuned from the small Aurora version, yet it approaches the state-of-the-art on global ensemble metrics and is competitive with the large version of Aurora. Our method is parameter and sample efficient, and generates skilful 15-day forecasts in 13 minutes. Our results demonstrate that deterministic foundation models can be efficiently extended into even stronger stochastic models.
Eliot Walt, Miltiadis Kofinas, Nikolaj Mücke +2
Vrije Universiteit Amsterdam · University of Amsterdam · Delft University of Technology +1
Limited-Area Models (LAMs) enable weather forecasting over regional domains at higher resolutions than what is computationally feasible for global models. At such high resolutions, machine learning approaches for weather prediction increasingly rely on ensemble methods to produce probabilistic forecasts. However, existing machine learning LAMs are not scalable due to relying on computationally costly diffusion models or inefficient graph neural networks. We tackle this by introducing a new hybrid CNN/GNN architecture, tailored to the LAM weather forecasting problem. Using this architecture, we construct the DET-LAM deterministic model, producing LAM forecasts both more efficiently and accurately than its graph-based competitor. We then tackle the ensemble forecasting problem, by using this architecture as a backbone for the generative model CRPS-LAM. CRPS-LAM is trained using a Continuous Ranked Probability Score (CRPS) objective, enabling efficient training and sampling in a single forward pass. This yields a speedup of ≈×39 compared to diffusion-based baselines. We evaluate our approach on regional domains in northern Europe, demonstrating that CRPS-LAM produces skillful and well-calibrated forecasts across a range of atmospheric variables.
Erik Larsson, Joel Oskarsson, Tomas Landelius +1
Linköping University · ETH AI Center · SMHI & Linköping University