Artificial intelligence pathways from weather to climate
Authors: Tom Beucler, J. David Neelin, Hui Su, Shivanshi Asthana, Chris Bretherton, Will Chapman, Costa Christopoulos, Spencer K. Clark, +5 more
Organizations: University of Lausanne, Lausanne, Switzerland. · University of California, Los Angeles, Los Angeles, CA, USA. · Hong Kong University of Science and Technology, Hong Kong SAR, China. · Allen Institute for AI, Seattle, WA, USA. · University of Colorado Boulder, Boulder, CO, USA. · California Institute of Technology, Pasadena, CA, USA. · NOAA/Geophysical Fluid Dynamics Laboratory, Princeton, NJ, USA. · Google Research, Mountain View, CA, USA. · New York University, New York, NY, USA.
Deep learning has made rapid advances in weather forecasting: autoregressive models trained on atmospheric reanalyses now rival dynamical models across nowcasting, medium-range, and subseasonal-to-seasonal lead times, producing well-calibrated ensemble forecasts at reduced cost. We review these advances and consider their extension to climate horizons, where the challenge shifts from initial-condition skill to producing reliable statistical responses under altered forcings. AI-powered climate prediction systems must produce credible forced responses to drivers (e.g., greenhouse gases, land-use change) typically outside the observed record. We propose two minimum requirements for AI in climate modeling: (i) external forcing agents must enter explicitly enough to support interventions in which they vary independently; and (ii) robustness must be stress-tested in out-of-distribution regimes, including extremes and counterfactual trajectories. Using leading AI autoregressive emulators and hybrid physics-AI models, we identify development and coupling challenges. Comparing the reported throughput of these models with that of GPU-ported dynamical models highlights how AI can reduce time-to-solution by advancing only the target variables at the required resolution and using longer time steps, rather than integrating a full high-frequency, multivariate state. Diverse AI downscaling strategies can partially substitute for explicit fine-scale resolution, paving the way toward inexpensive local hazard assessment across prediction horizons.
Figures & tables
Figure 1: AI forecast refinement for tropical cyclones, applied to Typhoon Saola (2023). (a) Track forecasts compared with IBTrACS observations. (b) Maximum wind speed and (c) minimum sea-level pressure forecasts. Black denotes observations; orange the ECMWF IFS ensemble mean; solid purple Pangu; yellow the GenCast ensemble mean; and dashed purple the Pangu post-processing ensemble mean. Shading in (b)–(c) shows the full memberwise range for Pangu and GenCast post-processing. (d) Near-real-time CCMP analysis and (e) deep-learning-downscaled wind magnitude at 06:00 UTC on 1 September 2023. Colors in (d)–(e) show wind speed in m s -1 , and black points indicate available in situ observations. Panels (a)–(c) use [ 49 ] ; panels (d)–(e) use [ 53 ] .
Figure 2: CliMA-S2S week-3 forecasts and probabilistic skill. Seven-day-mean 300-hPa geopotential height (top) and 2-m temperature (bottom) for lead days 15–21, initialized 9 July 1997. Left panels show the bias-corrected CliMA-S2S forecasts, middle panels show the corresponding ERA5 verification fields, and right panels show ranked probability skill score (RPSS) for climatological quintiles relative to ERA5 climatology. Because CliMA-S2S is run here as a single deterministic forecast, predictive uncertainty is estimated through lightweight leave-one-out statistical calibration rather than from an ensemble of perturbed initial conditions. Positive RPSS indicates improvement over climatology. Global RPSS and continuous ranked probability skill score (CRPSS) are indicated below the right panels.
Figure 3: Example formulation for including external forcing agents through restricted pathways, illustrated here for atmospheric tendencies. Land use and land cover (LULC) enter the subgrid/coupling term through surface properties and fluxes and the radiative term through albedo; aerosols enter the subgrid/coupling term through microphysics and the radiative term through direct and semi-direct effects; and greenhouse gases (GHGs) enter only through radiative forcing. Here, the direct atmospheric radiative tendency acts only on temperature, while other variables respond through subsequent dynamics, subgrid processes, and coupling. These pathways may become intertwined over multi-hour model steps, while persistent forcing modifies the evolving climate and its internal variability. Multiple realizations, denoted with superscript (i), for a given external forcing can generate large ensembles of coarse simulations, with optional downscaling; emissions-driven Earth system models must additionally represent the chemistry and biogeochemistry that determine GHG and aerosol concentrations.
Figure 4: Tropical precipitation variability in ACE2. Two-year Hovmöller diagrams of daily precipitation averaged from 10∘ S to 10∘ N show that a long ACE2 rollout trained on ERA5 reproduces statistical characteristics of ERA5 tropical precipitation, including eastward-propagating variability associated with the Madden–Julian Oscillation. Adapted from Fig. 6 of [ 107 ] .
Figure 5: Extreme precipitation in ACE2-SOM. Daily precipitation distributions from 50-year samples of reference slab-ocean-coupled ∼ 100 km and ∼ 400 km SHiELD simulations, and ACE2-SOM for present-day and unseen 3× CO 2 climates, regridded to a common 4 ∘ Gaussian grid. ACE2-SOM reproduces the target ∼ 100 km model’s extreme-precipitation frequency and its sensitivity to CO 2 -induced warming. Adapted from Fig. 6 of [ 115 ] .
Figure 6: AMIP constant CO 2 and slab-ocean abrupt 4xCO 2 experiments with ACE. Evolution of global annual mean 2-m temperature in simulations forced by observed sea-surface temperature and sea ice, but with CO 2 held constant at the 1979 level with the reference dynamical SHiELD model, ACE2-SHiELD, and ACE2S-SHiELD+ (a); and evolution of global daily and ensemble mean 2-m temperature in a 36-member ensemble of 90-day slab-ocean-coupled abrupt 4xCO 2 simulations with SHiELD, ACE2-SOM, and ACE2S-SHiELD+ (b). ACE2-SHiELD was trained solely on output from historical AMIP simulations; ACE2-SOM was trained solely on output from slab-ocean-coupled equilibrium-climate simulations with 1x, 2x, and 4x the present-day CO 2 ; and ACE2S-SHiELD+ was trained on a mixture of output from AMIP, equilibrium-climate, and ramped-SST, random-CO 2 simulations. Adapted from [ 116 ] .
Figure 7: Short-range loss and long-range climate error in ACE2. Training and validation losses for ACE2-ERA5 and ACE2-SHiELD are compared with an inference error computed from eight five-year simulations after each epoch. Short-range training loss and long-range climate error do not necessarily select the same model checkpoint. Adapted from Fig. S8 of [ 107 ] .
Figure 8: Coupled variability in SamudrACE. Five 200-year SamudrACE rollouts (colored lines) starting from initial conditions in the held-out period of the preindustrial CM4 simulation reproduce the approximate amplitude of Niño3.4 variability (a). Similarly, the ENSO precipitation response (b) and the seasonal sea-ice climatology (c) are approximately reproduced in a 40-year run during the hold-out period. However, SamudrACE exhibits a stronger preference for approximately three-year ENSO oscillations than CM4. Adapted from Figs. 3 and 4 of [ 124 ] .
Figure 9: Reported normalized throughput across GPU-accelerated dynamical, hybrid AI–physics, and AI emulator models. Each point represents one reported model configuration. Panel (a) plots SYPDnorm against nominal horizontal grid spacing Δx , whereas panel (b) accounts for the reported prognostic update time-interval through (Δx)2Δt ; both panels share the same ordinate scale. Black lines are descriptive weighted-least-squares fits in log–log space, assigning equal total weight to each model. Gray envelopes are nominal 95% prediction intervals.
Figure 10: Modeling choices for downscaling. Downscaling pathways connect scenario-forced global climate simulations to kilometer-scale simulations (center right) or observations and reanalyses (bottom right). The diagram distinguishes paired mappings, such as emulation and super-resolution, from unpaired bias correction across the model (top) and observational (bottom) domains; horizontal position indicates approximate horizontal grid spacing.
Atmospheric predictability declines rapidly beyond the next ten days, such that forecasts at longer lead times primarily convey large-scale trends rather than specific states. Yet in a warming world, improving early warnings of extreme heat is an increasingly critical challenge. Here we evaluate six state-of-the-art deep learning weather emulators - Pangu-Weather, FuXi, ArchesWeather, AIFS, GraphCast and Aurora - alongside leading dynamical systems and statistical baselines in forecasting global near-surface temperature and extreme heat at lead times of 10-15 days. We find that several emulators rival or even surpass physics-based forecasts in deterministic temperature skill, but do so at the cost of reduced spectral fidelity, in a process widely known as blurring. While all models show some degree of predictive skill for extreme heat, most emulators under-represent peak intensities, and IFS recall is greater than that of any of the emulators. These results highlight both the emerging potential of AI to enhance extended range temperature prediction, and the remaining challenges in delivering reliable, actionable early warnings in a changing climate.
Cas Decancq, Thomas Mortier, Jessica Keune +1
1Hydro-Climate Extremes Lab (H-CEL), Department of Environment, Ghent University, Ghent, Belgium. · 2European Centre for Medium-Range Weather Forecasts (ECMWF), Reading, United Kingdom.
Kilometer-scale convection shapes precipitation extremes, tropical organization, and cloud feedbacks, but most global atmospheric models approximate these processes at 25-100 km resolution. Global storm-resolving physics models resolve convective systems explicitly, but at a cost -- roughly one MWh per simulated day on exascale supercomputers -- that limits long-duration simulation. We introduce STRATA (Storm-resolving Tile-based autoRegressive Atmosphere Transformer Architecture), the first autoregressive AI emulator for global storm-resolving atmospheric dynamics. STRATA is trained on the highest-resolution atmospheric dataset yet used for global AI emulation: 17 days of SCREAM physics-model output at 4.9-km resolution (~25 million grid cells) sampled every 10 minutes. Our central premise is that on 10-minute timescales atmospheric dynamics are predominantly local, so training on small spatial tiles trades scarce global temporal samples for abundant local spatial samples and enables global rollout via overlapping-tile blending. STRATA combines 3D patch embedding and local 3D neighborhood attention, a novel Stereographic Rotary Position Embedding (StereoRoPE) for grid-invariant encoding, and a pixel-space de-aliasing decoder that suppresses patch-scale rollout artifacts. An iso-FLOP scaling study reveals that km-scale emulation requires ~10x more FLOPs per grid point than coarse-resolution AI weather models, consistent with the higher information density of convective-scale dynamics. Trained on only 17 days of data, STRATA produces stable 24-hour global rollouts with realistic km-scale dynamics across diverse regimes, though large-scale biases develop with lead time. It achieves 48 simulation days per megawatt-hour -- about 50 times better energy efficiency than the SCREAM physics model -- and 741 simulated days per wall-clock day at 512 H100 GPUs. Code and dataset are publicly available.
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best physics-based models and they are exclusively initialized with and trained on analysis data. As a result, they cannot directly make use of observations, and any biases in the analysis are inherited by the forecast. WeatherNext 3 addresses these shortcomings and establishes a new state-of-the-art for probabilistic medium-range forecasting skill. First, WeatherNext 3 generates new forecasts every hour (rather than every 6 hours like traditional global models) by ingesting low-latency geostationary satellite data. Second, WeatherNext 3's temporal and spatial resolution are on par with physics-based global models, with hourly time steps and 0.1 degree resolution for single-level variables, including solar radiation and cloud cover. Third, WeatherNext 3 moves beyond traditional analysis variables by learning to predict satellite-derived precipitation estimates, as well as tropical cyclone and station observations. Modelling sparse station data allows WeatherNext 3 to make 2m temperature and dewpoint predictions at any location and time, conditioned on local geographical features, with substantially lower error than competing global models, even when evaluated against unseen stations. Together, WeatherNext 3's capabilities move operational AI-based weather forecasting beyond emulating the traditionally distinct stages of data assimilation, forecasting and post-processing, which helps to further push the frontier of performance and granularity for global weather prediction.