Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instantiate this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill. Experiments on SEVIR and MRMS benchmarks show that NowcastDiT achieves state-of-the-art performance in both perceptual quality and meteorological skill. These results suggest that standard DiT can serve as an effective foundation for precipitation nowcasting.
Figures & tables
Figure 1: Comparison of diffusion modeling paradigms for precipitation nowcasting. (a): Specialized diffusion nowcasters encode domain knowledge through task-specific models or losses. (b): A standard DiT is applicable for precipitation nowcasting. (c): NowcastDiT adds minimal precipitation-specific adaptation towards a standard DiT without changing its backbone.
Figure 2: Overview of NowcastDiT . Left: A standard Diffusion Transformer predicts future precipitation conditioned on encoded observations, with QK-Norm and 3D RoPE. Top right: The dynamics-aware noise prior (DyPro) schedules cross-frame noise correlation from temporally correlated at the noise endpoint to independent at the clean endpoint. Bottom right: Timestep-aware post-training shifts the reward emphasis from structure to detail along the denoising trajectory.
SEVIR
MRMS
Method
CSI ↑
CSI 181↑
CSI 219↑
HSS ↑
LPIPS ↓
SSIM ↑
Method
CSI ↑
CSI 16↑
CSI 32↑
HSS ↑
LPIPS ↓
SSIM ↑
ConvLSTM
0.2912
0.0684
0.0336
0.3571
0.2989
0.7257
ConvLSTM
0.2159
0.0514
0.0172
0.2877
0.3013
0.8978
PhyDNet
0.2874
0.0664
0.0211
0.3523
0.3081
0.7252
PhyDNet
0.2332
0.0635
0.0268
0.3111
0.3005
0.9001
Earthformer
0.2752
0.0496
0.0213
0.3391
0.3376
0.7142
Earthformer
0.2353
0.0691
0.0297
0.3152
0.3001
0.8994
SimVP
0.2951
0.0753
0.0416
0.3630
0.3113
0.7247
SimVP
0.2366
0.0657
0.0277
0.3155
0.2992
0.9006
AlphaPre
0.2980
0.0860
0.0436
0.3677
0.2897
0.7301
AlphaPre
0.2309
0.0587
0.0211
0.2933
0.3014
0.8944
Table 1: Performance comparison on SEVIR and MRMS. ↑ indicates higher is better and ↓ indicates lower is better. Bold and underlined values denote the best and second-best results, respectively.
Figure 3: Qualitative comparison on SEVIR. Forecasts are shown every 10 minutes up to 100 minutes. Colors indicate VIL values.
Figure 4: Qualitative comparison on MRMS. Forecasts are shown every 20 minutes up to 200 minutes. Colors indicate precipitation rate in mm h -1 .
Figure 5: Analysis on SEVIR. Bar lengths indicate CSI gains relative to the first setting in each panel; labels report absolute CSI values.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
SEVIR
MRMS
Architecture
Parameters
82M
Input shape
25×1×128×128
24×1×256×256
Encoder stages
4
Residual blocks / encoder stage
2
Encoder channels
[128, 256, 512, 512]
Appendix
Table 2: VAE architecture and training configuration. Parameters include the encoder, decoder, and latent projections.
Setting
SEVIR
MRMS
Architecture
Parameters
130M
Frames (history / future)
5 / 20
4 / 20
Transformer blocks
12
Hidden dimension
768
Feedforward dimension
3072
Appendix
Table 3: Transformer architecture, flow-matching pretraining, and inference configuration. Parameters include the conditioner, time embedding, and output layer. Inference settings refer to the pretrained-model evaluation configurations.
Setting
SEVIR
MRMS
Optimization
Optimizer
AdamW, (β1,β2)=(0.9,0.95)
Learning rate
10−5 , constant
Batch size
8
Policy ratio clipping range εc
10−4
Advantage clipping
[−5,5]
Appendix
Table 4: Post-training, reward, and evaluation configurations. Reward thresholds are expressed in VIL values for SEVIR and mmh−1 for MRMS.
Setting
SEVIR
MRMS
Data variable
VIL
Precipitation rate
Spatial / temporal res.
1km / 5min
0.01∘ / 10min
Forecast frames
5→20
4→20
Frame size
128×128
256×256
Train / validation / test
59,530 / 28,145 / 7,220
6,807,528 / 12,000 / 12,000
Appendix
Table 5: Dataset configurations. Spatial resolution refers to the source radar grids, and frame size denotes the model input and output size. Forecast frames are shown as observed → predicted; split sizes count sequences.
Method
CSI ↑
CSI-181 ↑
CSI-219 ↑
HSS ↑
LPIPS ↓
SSIM ↑
DiT
0.2926
0.0948
0.0555
0.3733
0.1545
0.7048
+ 3D RoPE, QK-Norm
0.3012
0.1082
0.0676
0.3647
0.1317
0.6812
+ DyPro
0.3195
0.1204
0.0752
0.4092
0.1486
0.7099
+ RL
0.3240
0.1241
0.0750
0.4148
0.1469
0.7197
Appendix
Table 6: Cumulative component ablation on SEVIR. ↑ indicates higher is better and ↓ indicates lower is better. Bold and underlined values denote the best and second-best results, respectively.
RL
Guidance
CSI ↑
CSI-181 ↑
CSI-219 ↑
HSS ↑
LPIPS ↓
SSIM ↑
✘
✘
0.2934
0.0956
0.0560
0.3749
0.1541
0.7041
✘
✔
0.3195
0.1204
0.0752
0.4092
0.1486
0.7099
✔
✘
0.3109
0.1050
0.0588
0.3961
0.1580
0.7156
✔
✔
0.3240
0.1241
0.0750
0.4148
0.1469
0.7197
Appendix
Table 7: Effects of RL and classifier-free guidance on SEVIR. ↑ indicates higher is better and ↓ indicates lower is better. Bold and underlined values denote the best and second-best results, respectively.
α
CSI ↑
CSI-181 ↑
CSI-219 ↑
HSS ↑
LPIPS ↓
SSIM ↑
0.00
0.3012
0.1082
0.0676
0.3647
0.1317
0.6812
0.25
0.3188
0.1184
0.0759
0.4075
0.1485
0.7090
0.33
0.3191
0.1192
0.0742
0.4080
0.1483
0.7097
0.50
0.3195
0.1204
0.0752
0.4092
0.1486
0.7099
0.66
0.3185
0.1194
0.0740
0.4077
0.1495
0.7098
0.75
0.3184
0.1184
0.0750
0.4084
0.1494
0.7092
Appendix
Table 8: DyPro correlation strength α on SEVIR. α=0 corresponds to the standard i.i.d. Gaussian noise prior. ↑ indicates higher is better and ↓ indicates lower is better. Bold and underlined values denote the best and second-best results, respectively.
Reward
Static
Timestep-aware
Mixed
(Ms+Md)/2
(1−t)Ms+tMd
CSI
(CL+CH)/2
(1−t)CL+tCH
Perceptual
(0.1S−0.9P)/2
0.1(1−t)S−0.9tP
Appendix
Table 9: Terminal rewards for the six ablation settings.
Figure 6: Forecast skill across lead times on MRMS and SEVIR. The horizontal axes show forecast lead time in minutes; higher CSI and HSS indicate better performance.
Figure 7: Additional reward ablations on SEVIR. LPIPS (left) and RMSE (right) during RL training for the rewards in Table 9 . The legend follows Figure 5 (d); the timestep-aware perceptual reward uses 0.1(1−t)SSIM−0.9tLPIPS . Lower values indicate better performance for both metrics.
Figure 8: Qualitative comparison on SEVIR (case 802). Five input observations span −20 to 0 minutes at 5-minute intervals. Ground truth and forecasts are shown every 10 minutes up to 100 minutes. Colors indicate VIL values.
Figure 9: Qualitative comparison on SEVIR (case 6155). Five input observations span −20 to 0 minutes at 5-minute intervals. Ground truth and forecasts are shown every 10 minutes up to 100 minutes. Colors indicate VIL values.
Figure 10: Qualitative comparison on MRMS (case 239). Four input observations span −30 to 0 minutes at 10-minute intervals. Ground truth and forecasts are shown every 20 minutes up to 200 minutes. Colors indicate precipitation rate in mm h -1 .
Figure 11: Qualitative comparison on MRMS (case 1149). Four input observations span −30 to 0 minutes at 10-minute intervals. Ground truth and forecasts are shown every 20 minutes up to 200 minutes. Colors indicate precipitation rate in mm h -1 .
School of Intelligence Science and Technology, Nanjing University, Suzhou, 215163, China. · State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, 210023, China.