Long-term precipitation nowcasting requires modeling radar-echo evolution while preserving localized high-intensity structures. Recent radar-specific studies motivate location-aware prediction and separating echo displacement from intensity change. However existing encoders learn historical representations mainly from final forecast errors. We propose PrecipJEPA, which couples a structured forecasting path with an auxiliary path that enriches its encoder from observed radar history. In the forecasting path, an online encoder first converts the observations into spatiotemporal tokens. The Task-Driven Future-State Predictor (TFP) combines these tokens with a recent-dynamics summary and spatiotemporal queries to construct future radar states. The Parallel Motion-Source Renderer (PMSR) decodes these states into motion and source-sink fields that transform the latest observation into future frames. During joint training, the History-Masked JEPA (H-JEPA) operates on the auxiliary path to predict masked historical features from visible context, directly supervising the same online encoder from the observed sequence. Experiments on SEVIR and MeteoNet show that PrecipJEPA improves highest-threshold CSI by 118.6% and 35.1%, respectively, over the strongest baselines, while maintaining the highest mean CSI throughout the 3-hour forecast.
Figures & tables
Figure 1: Overview of PrecipJEPA. TFP maps historical representations to lead-specific future states, which PMSR renders into radar forecasts, while H-JEPA supervises the shared encoder through masked-latent prediction during training.
SEVIR
MeteoNet
Method
CSI ↑
CSI-p4 ↑
CSI-p16 ↑
HSS ↑
LPIPS ↓
CSI-last ↑
CSI 181 ↑
CSI 219 ↑
CSI ↑
CSI-p4 ↑
CSI-p16 ↑
HSS ↑
LPIPS ↓
CSI-last ↑
CSI 24 ↑
CSI 32 ↑
SimVP [ 4 ]
0.2073
0.2277
0.2405
0.2518
0.4348
0.1362
0.0276
0.0147
0.2598
0.3251
0.3562
0.3589
0.3444
0.1290
0.1974
0.0674
Earthformer [ 6 ]
0.2006
0.2193
0.2384
0.2456
0.4434
0.1366
0.0168
0.0024
0.2398
0.2968
0.3393
0.3370
0.4130
0.1309
0.1924
0.0362
AlphaPre [ 10 ]
0.2162
0.2381
0.2495
0.2652
0.4192
0.1432
0.0311
0.0139
0.2708
0.3322
0.3624
0.3723
0.3308
0.1366
0.2191
0.0771
exPreCast [ 18 ]
0.2142
0.2365
0.2497
0.2726
0.3617
0.1472
0.0391
0.0161
0.2583
0.3230
0.3668
0.3659
0.3170
0.1532
0.2279
0.1043
CasCast [ 7 ]
0.1977
0.2105
0.2333
0.2549
0.4108
0.1341
0.0337
0.0145
0.2191
0.2299
0.2224
0.3131
0.3454
0.1596
0.2053
0.0824
Table 1: Quantitative comparison on SEVIR and MeteoNet. Aggregate metrics cover all thresholds and forecast times; CSI-last covers the final forecast hour. CSI 181 , CSI 219 , CSI 24 , and CSI 32 report the two highest standard thresholds of each dataset. Best results are in bold and second-best results are underlined .
Figure 2: Long-term evaluation on SEVIR: (a) mean CSI over the full 180-minute forecast; (b) example forecasts from different models.
Variant
TFP
PMSR
H-J
CSI ↑
HSS ↑
CSI-last ↑
CSI 219 ↑
HSS 219 ↑
M0 (Base)
–
–
–
0.2082
0.2658
0.1477
0.0130
0.0239
M1 (w/o TFP)
–
✓
✓
0.2162
0.2761
0.1344
0.0305
0.0528
M2 (w/o PMSR)
✓
–
✓
0.2082
0.2662
0.1441
0.0135
0.0247
M3 (w/o H-J)
✓
✓
–
0.2215
0.2841
0.1374
0.0342
0.0592
PrecipJEPA
✓
✓
✓
0.2315
0.2974
0.1493
0.0376
0.0657
Table 2: Component ablation on SEVIR and MeteoNet. H-J denotes H-JEPA. In M0 and M1, TFP is replaced by the basic future-state predictor; a dash under PMSR indicates use of the residual head. Best results are in bold .
Figure 3: H-JEPA training strategies on (a) SEVIR and (b) MeteoNet. M3 trains without H-JEPA, M4 uses it for encoder pretraining before forecast-only training, and PrecipJEPA jointly trains H-JEPA with the forecasting path.
Precipitation nowcasting predicts the spatiotemporal evolution of future radar echoes from historical radar echo sequences, thereby estimating the occurrence, development, and movement of precipitation over the near term. In recent years, deep learning has become an important approach to precipitation nowcasting. Although state-of-the-art models can generally capture the overall spatial distribution of future precipitation, their predictions still exhibit substantial biases in radar echo intensity at individual locations. This observation motivates a more targeted strategy for reducing forecast errors. Instead of regenerating an entire radar echo sequence without spatial constraints, the predicted precipitation structure can be used to guide the refinement of echo intensities at individual locations. This structure-guided refinement directly targets echo intensity biases. Accordingly, we propose FreCast, a two-stage framework for radar echo prediction. The first stage generates an initial forecast of future radar echoes. The second stage uses the spatial structure of the initial forecast as a constraint to further correct intensity biases at individual locations in the first-stage prediction. Experiments on three datasets demonstrate that FreCast achieves consistent improvements across forecast skill metrics. Qualitative results further show that FreCast better preserves rainband continuity and intense precipitation structures at longer lead times.
Heping Fang, Zihuai Yin, Kaicheng Mao +2
Department of Statistics and Data Science, Southern University of Science and Technology, Shenzhen 518055, China · Department of Computer Science and Engineering, Southern University of Science and Technology, Shenzhen 518055, China · Guangdong Provincial Key Laboratory of Brain-Inspired Intelligent Computation, Department of Computer Science and Engineering and the Department of Statistics and Data Science, Southern University of Science and Technology, Shenzhen 518055, China
Precipitation nowcasting is increasingly being approached with deep learning models that learn directly from recent radar observations. Although such models can efficiently capture short-term precipitation motion, they often lack broader contextual information about the meteorological conditions under which rainfall develops. This paper investigates whether lightweight temporal context can improve radar-based nowcasting, particularly for high-intensity rainfall. We propose the Time-Aware Small-Attention U-Net (TA-SmaAt-UNet), which extends the core SmaAt-UNet model with temporal conditioning layers that use cyclical encodings of time-of-day and time-of-year to modulate intermediate feature representations. Experiments on KNMI radar precipitation data show that temporal conditioning is most beneficial for rare, high-intensity precipitation events, while also improving the representation of seasonal variability and predicted rainfall-intensity distributions. A layer conductance analysis further indicates that the added temporal conditioning layers are actively used by the model despite their small parameter cost. These findings suggest that simple, physically motivated temporal context can improve the realism and reliability of deep learning-based precipitation nowcasts. The implementation of our models and training setup is available on \href{https://github.com/gijsvn/TA-SmaAt-UNet}{GitHub}.
Gijs van Nieuwkoop, Siamak Mehrkanoon
Department of Information and Computing Sciences, Utrecht University, Utrecht, The Netherlands
Precipitation nowcasting is a vital spatio-temporal prediction task for meteorological applications but faces challenges due to the chaotic property of precipitation systems. Existing methods predominantly rely on single-source radar data to build either deterministic or probabilistic models for extrapolation. However, the single deterministic model suffers from blurring due to MSE convergence. The single probabilistic model, typically represented by diffusion models, can generate fine details but suffers from spurious artifacts that compromise accuracy and computational inefficiency. To address these challenges, this paper proposes a novel coarse-to-fine Vision Mamba Unet and residual Diffusion (VMU-Diff) based precipitation nowcasting framework. It realizes precipitation nowcasting through a two-stage process, i.e., a deterministic model-based coarse stage to predict global motion trends and a probabilistic model-based fine stage to generate fine prediction details. In the coarse prediction stage, rather than single-source radar data, both radar and multi-band satellite data are taken as input. A spatial-temporal attention block and several Vision mamba state-space blocks realize multi-source data fusion, and predict the future echo global dynamics. The fine-grained stage is realized by a spatio-temporal refine generator based on residual conditional diffusion models. It first obtains spatio-temporal residual features based on coarse prediction and ground truth, and further reconstructs the residual via conditional Mamba state-space module. Experiments on Jiangsu SWAN datasets demonstrate the improvements of our method over state-of-the-art methods, particularly in short-term forecasts.
Chunlei Shi, Hao Li, Yufeng Zhu +6
Department of Automation, Southeast University, Nanjing 210096, China