Just Weather Scoring: Efficient End-to-end Nowcasting with Distributional Diffusion
Abstract
Generative diffusion models are well-suited for probabilistic precipitation nowcasting, but existing approaches often rely on separately trained compression or deterministic forecasting components and remain costly at inference due to iterative denoising. We introduce Just Weather Scoring (JWS), a single-stage, end-to-end diffusion model which addresses both issues by forecasting directly in radar space and enabling few-step generation. Radar-space modeling greatly simplifies training and inference and eliminates uncertainty arising from lossy compression. JWS combines Masked Asynchronous Diffusion, a timestep-sampling scheme that preserves clean context while adapting diffusion training to high-dimensional spatio-temporal data, with a simple scoring-rule objective that aligns training with probabilistic forecasting and unlocks few-step generation. On the SEVIR and MeteoNet benchmarks, JWS achieves state-of-the-art probabilistic forecasting performance at reduced training and inference cost. Even our smallest model remains competitive using substantially fewer parameters and more than 17x faster inference.
Figures & tables
| Method | Stages | Param | CRPS | SSIM | HSS | CSI | RI ⋆ | |
| Deterministic | ConvLSTM ( Shi et al., 2015 ) (NeurIPS ’15) | 1 | 14M | 0.0264 | 0.7749 | 0.5232 | 0.4102 | – |
| PredRNN ( Wang et al., 2017 ) (TPAMI ’22) | 1 | 47M | 0.0271 | 0.7497 | 0.5192 | 0.4045 | – | |
| PhyDNet ( Guen & Thome, 2020 ) (CVPR ’20) | 1 | 14M | 0.0253 | 0.7649 | 0.5311 | 0.4198 | – | |
| SimVP ( Gao et al., 2022a ) (CVPR ’22) | 1 | 16M | 0.0259 | 0.7772 | 0.5280 | 0.4153 | – | |
| EarthFormer ( Gao et al., 2022b ) (NeurIPS ’22) | 1 | 9M | 0.0251 | 0.7756 | 0.5411 | 0.4310 | – | |
| Generative | NowcastNet ( Zhang et al., 2023 ) (Nature ’23) | 1 | 35M | 0.0283 | 0.5696 | 0.5365 | 0.4152 | – |
| Method | Stages | Param | CRPS | SSIM | HSS | CSI | RI ⋆ | |
| D. | EarthFormer ( Gao et al., 2022b ) (NeurIPS ’22) | 1 | 9M | 0.0224 | – | – | 0.2831 | – |
| Generative | NowcastNet ( Zhang et al., 2023 ) (Nature ’23) | 1 | 35M | 0.0277 | – | – | 0.2955 | – |
| PreDiff ( Gao et al., 2023 ) (NeurIPS ’23) | 2 | 105M | 0.0197 | – | – | 0.2546 | – | |
| CasCast ( Gong et al., 2024 ) (ICML ’24) | 3 | 402M | 0.0180 | – | – | 0.3156 | – | |
| FREUD ( Schusterbauer et al., 2026b ) (CVPR ’26) | 2 | 521M | 0.0193 | 0.7312 | 0.2082 | 0.1417 | – | |
| JWS-T/32 ( ours ) | 1 | 9M | 0.0143 | 0.7887 | 0.4026 | 0.2851 | 0.2376 | |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | JWS-T | JWS-S | JWS-B | JWS-L |
| Parameters (M) | 9 | 24 | 94 | 340 |
| High-resolution layers (down & up) | 2 | 2 | 2 | 2 |
| Bottleneck | 6 | 8 | 8 | 20 |
| High resolution hidden dimension | 192 | 256 | 512 | 768 |
| Bottleneck hidden dimension | 256 | 384 | 768 | 1024 |
| Attention heads (HiRes/Bottleneck) | 3/4 | 4/6 | 8/12 | 12/16 |
| Num. Samples | CRPS | SSIM | HSS | CSI | RI |
| 64 | |||||
| 128 | |||||
| 256 | |||||
| 512 | |||||
| 1024 |