Kilometer-scale regional weather forecasting is essential for local weather warnings and weather-sensitive decisions. Existing data-driven approaches often rely on numerical forecasts for large-scale guidance or require additional training of global forecasting components. Pretrained global weather models offer an efficient source of large-scale forecasts, motivating their reuse to guide high-resolution regional prediction. However, this coupling requires aligning global and regional representations across different grids and integrating global guidance with local interactions to advance regional states. We propose ScaleCast, a regional forecasting framework that addresses these challenges through Global-Regional Alignment. Its Global-Regional Conversion module aligns joint global and regional representations with regional locations, while the Global-Regional Alignment and Dynamics block combines aligned guidance with regional neighborhood interactions. Experiments using ERA5 global analyses on a 0.25-degree grid and CERRA regional reanalysis at 5.5 km spacing demonstrate improved regional forecasts across surface and upper-air variables, with a single trained model supporting multiple global forecast drivers (i.e., Pangu-Weather, GraphCast, and HRES) without specific retraining. Fine-tuning on HRRR at 3 km spacing further demonstrates the framework's adaptability to a different regional domain and spatial resolution. Windstorm case studies show improved cyclone positioning and core-pressure estimates, while comparisons with HadISD station observations show closer agreement with local temperature and humidity changes.
Figures & tables
Figure 1: Coarse global and finer regional projected grids. The two grids sample the same atmospheric system at different locations and spatial resolutions. The surrounding gray shading denotes the boundary region. Background colors indicate mean sea level pressure (MSLP).
Figure 2: ScaleCast framework during training. (a) Training pipeline from global analysis at t+3 h, the regional state at t , and terrain to the regional prediction at t+3 h through feature encoding, stacked GRAD blocks, and increment decoding. (b) A Global–Regional Alignment and Dynamics (GRAD) block from (a), combining GRConv alignment, regional attention, and gated updates. (c) Details of cross-grid token alignment and windowed attention within (b).
Surface RMSE ↓
Upper-air RMSE ↓
Method
MSLP
T2m
RH2m
T500
T850
Z500
Z850
RH500
RH850
ECMWF-HRES
92.9
1.331
–
0.733
0.828
0.772
0.711
22.680
12.622
CERRA NWP
97.1
0.975
6.732
0.793
0.744
0.796
0.750
17.596
10.668
Pangu
93.9
1.461
–
0.725
0.840
0.775
0.685
–
–
GraphCast
91.8
1.431
–
0.695
0.784
0.733
0.671
–
–
ML-LAM
95.1
1.133
7.347
0.725
0.844
0.851
0.694
17.416
10.908
Table 1: Recursive regional forecast errors through +30 h. Pixel-pooled RMSE over ten 3-hour leads in 2022. Lower is better.
Figure 3: Regional errors for two cases: T2m at +9 h from 16 November 2022, 00 UTC (left), and RH2m at +27 h from 5 April 2022, 00 UTC (right). CERRA shows absolute values; forecasts show deviations from CERRA. Pangu and GraphCast RH2m are unavailable and omitted.
Design choices
Surface RMSE
Upper-air RMSE
Type
Module
MSLP ↓ (Pa)
T2m ↓ (K)
RH2m ↓ (pp)
T500 ↓ (K)
T850 ↓ (K)
Z500 ↓ (dam)
Z850 ↓ (dam)
RH500 ↓ (pp)
RH850 ↓ (pp)
nMSE ↓ ( 10−2 )
Attention
Differential
41.1
0.724
5.203
0.379
0.459
0.477
0.341
7.489
6.105
2.733
Gated
36.7
0.705
5.226
0.338
0.427
0.429
0.309
7.420
5.931
2.713
Cross-attention
35.9
0.708
5.219
0.337
0.430
0.410
0.298
7.623
6.020
2.765
Our attention
34.6
0.710
5.233
0.346
0.428
0.409
0.291
7.396
5.893
2.702
Spatial
Scale + terrain
62.7
0.931
5.484
0.510
0.541
0.657
0.512
9.337
6.652
3.579
Table 2: Module ablations. Bold and underlined values indicate the best and second-best results.
Figure 4: Windstorm position and core-pressure diagnostics. (a–d) MSLP at +72 h and cyclone tracks. (e) Position errors for event 1804. (f) Time-mean absolute core-pressure bias relative to CERRA within 200 km of Hodges centres.
Figure 5: T2m and RH2m against HadISD stations. (a) Station locations and terrain. (b,c) T2m and RH2m forecasts against HadISD observations through +72 h at S1(top) and S2(bottom).
Figure 6: Cross-dataset experiment results. (a) HRRR analysis of 2 m temperature. (b–d) Models’ forecasts minus the same analysis. (e) RMSE comparison for T2m and MSLMA.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Variable
Symbol
Level / height
Unit
Channels
Global atmospheric inputs 24 channels
2-m temperature
T2m
2 m
K
1
Mean sea-level pressure
MSLP
Sea level
Pa
1
10-m zonal wind
u10m
10 m
ms−1
1
10-m meridional wind
v10m
10 m
ms−1
1
Temperature
T
50, 500, 850, 1000
K
4
Appendix
Table 3: Atmospheric variables used by ScaleCast. Global atmospheric inputs contain 24 channels, while the regional state and prediction contain nine channels. Pressure levels are given in hPa, and physical units are listed before normalization.
Figure 7: Three-driver frozen-weight comparison. Per-lead pixel-pooled RMSE for the same regional checkpoint driven by Pangu, GraphCast, or HRES.
Variable
Unit
ScaleCast + Pangu ↓
ScaleCast + GFS ↓
Δ (%)
MSLP
Pa
89.9
171.3
+90.6
T2m
K
0.958
1.732
+80.9
RH2m
pp
6.737
8.727
+29.5
T500
K
0.700
1.088
+55.5
T850
K
0.748
1.043
+39.4
Z500
dam
0.735
1.395
+89.9
Appendix
Table 4: Frozen-driver transfer through +30 h. Pixel-pooled RMSE over ten 3-hour lead times. Δ is the relative change from Pangu to GFS forcing; lower is better.
Figure 8: Lead-dependent driver-shift test. RMSE of the frozen ScaleCast checkpoint under fixed-initialization Pangu and GFS forcing. CERRA analyses are used only for verification.
Figure 9: Complete recursive forecast trajectories through +30 h. Pixel-pooled RMSE for the completed regional forecasting models, together with the same-cycle CERRA numerical forecast.
Figure 10: Mean absolute error through +30 h. Per-lead pixel-pooled MAE for the methods in Table 1 . Solid curves highlight the frozen ScaleCast driver variants; dashed and dotted curves show external, global, and regional comparison systems. Missing curves denote diagnostics unavailable from the source forecast.
Figure 11: Absolute fields for two cases: T2m from 16 November 2022, 00 UTC at +9 h (left), and RH2m from 5 April 2022, 00 UTC at +27 h (right). Each variable uses a shared colour scale across the reference and forecasts, in ∘ C and %, respectively. The square window and post-hoc case selection match Figure 3 .
Figure 12: November case and storm-relative displacement. Event 1812, initialized at 00 UTC on 16 November 2022. a–d, MSLP at +72 h and cumulative Hodges and independently tracked field centres. e, Along- and cross-track displacement at eight shared fixes with equal kilometre scales; positive values indicate ahead of and right of reference motion. Large symbols mark means. f, Mean distance and mean absolute displacement components.
Figure 13: Archived four-event diagnostics. Paired errors use shared times ( n ); coverage uses all eligible fixes. Current event-1804 position results appear in Figure 4 .
Figure 14: Archived four-event lead curves. Left: position error. Right: signed central-pressure error. Current event-1804 position results appear in Figure 4 .
Figure 15: Archived February tracking diagnostics. Initialization: 17 February 2022, 00 UTC. Maps show MSLP at +72 h and cumulative Hodges and field-derived tracks; the right panels show position errors for events 1804 and 1805. The event-1804 curve predates the current comparison in Figure 4 . The shaded +66 h point is matched only by SC + GraphCast.
Figure 16: Spatial temperature and humidity errors against HadISD. Initialization: 17 February 2022, 00 UTC; lead: +9 h. a–f, T2m errors at 1,626 common stations; g–l, RH2m errors at 1,140 common stations. Colours show forecast minus observed values.
Design choices
Surface RMSE
Upper-air RMSE
Module
Training
MSLP ↓ (Pa)
T2m ↓ (K)
RH2m ↓ (pp)
T500 ↓ (K)
T850 ↓ (K)
Z500 ↓ (dam)
Z850 ↓ (dam)
RH500 ↓ (pp)
RH850 ↓ (pp)
nMSE ↓ ( 10−2 )
Spatial training strategies
Scale + terrain
Grouped
40.8
0.766
5.336
0.383
0.462
0.476
0.336
8.173
6.217
3.009
Scale + terrain
MS + grad.
43.4
0.779
5.366
0.383
0.464
0.500
0.357
8.189
6.225
3.038
Scale + terrain 10
FAMO + noise
55.8
0.782
5.253
0.428
0.489
0.626
0.469
9.009
6.644
3.305
Humidity objectives
Appendix
Table 5: Additional training comparisons. RMSE and normalized MSE on 2,919 validation samples from 2023 at +3 h, conditioned on target-time ERA5 analyses. Lower is better. Bold and underlined values indicate the best and second-best results within each block.
Figure 17: Coarse and regional grids in different coordinate systems. (a,b) Actual footprints of the 128×128 and 128×208 coarse grids, with 0.25∘ angular spacing, overlaid on the 512×512 CERRA grid. Expansion covers the regional side areas outside the original coarse bounds; the dashed outline marks the original footprint. (c) The same regional grid in relative native coordinates, with approximately 5.5 km spacing and geographic coordinate contours. Grid lines are subsampled.
Figure 18: HRRR analysis fields underlying the regional comparisons. Rows show 09 UTC on 2 and 28 December 2022; columns show T2m and MSLMA.
Department of Earth, Environmental, and Atmospheric Sciences, Western Kentucky University, Bowling Green, KY, USA · NASA Goddard Space Flight Center, Greenbelt, MD, USA · The University of Texas at Austin, Austin, TX, USA +3
Swiss Data Science Center, ETH, Zürich, Switzerland · Expertise Center for Climate Extremes, University of Lausanne, Lausanne, Switzerland · Faculty of Geosciences and Environment, University of Lausanne, Lausanne, Switzerland +1