While AI weather models now rival operational forecasts, how they represent the atmosphere internally remains an open question: feature attribution reveals which input patterns matter, not what the model computes or how it combines information internally. We train sparse autoencoders (SAEs) on GraphCast to uncover its learned concepts, using atmospheric rivers as our phenomenon of focus. Both standard and Matryoshka SAEs show GraphCast computes atmospheric river intensity, measured by integrated vapor transport (IVT), as a stable internal variable, despite IVT being neither an input nor a target. In contrast to the unstructured concept retrieval of the standard SAE, the Matryoshka SAE orders concepts by importance and exposes their relations. Atmospheric river concepts persist across depth and direct interventions confirm causality. This method offers a way to find internal variables and determine which of them the model actually relies on, which is a prerequisite for asking whether those variables remain meaningful as the phenomenon changes under a warming climate.
Figures & tables
Figure 1: Top: GraphCast maps ERA5 grid inputs at times t−1 and t ( 227 variables per node) onto an icosahedral mesh, where a 16 -layer processor updates the 512 -dimensional mesh-node activations before the decoder returns the forecast at t+1 . Bottom: We train SAEs post-hoc on the mesh-node activations at layers 0 , 8 , and 15 , each encoding the 512 -dimensional activation into a wider ( R4096 ), sparsely active dictionary and reconstructing it. Inset: One of many AR concepts. The activation pattern of Concept 1592 at layer 8 , which responds to strong IVT over western North America.
Figure 2: (A) Concept 99 tracks regional AR intensity in four regions among the most frequently affected by ARs [ 22 , 23 , 24 ] : western North America ( 30 – 50∘ N, 115 – 130∘ W), western Europe ( 37 – 57∘ N, 8∘ W– 7∘ E), western South America ( 30 – 50∘ S, 62 – 77∘ W), and eastern Australia ( 22 – 42∘ S, 147 – 162∘ E). Displayed are region-mean activation of Matryoshka-L8 Concept 99 (red, left axis) against region-maximum IVT (blue, right axis) over AR timesteps, averaged into monthly means. Each panel displays the four-year window with the strongest monthly correlation. Full training-period correlation is printed above each panel. (B) Every mesh node where Concept 99 and child Concept 3153 fire across 8,000 sampled timesteps, colored by the mean IVT at that node over the timesteps it fires; blank regions are nodes where the concept never fires. (C) Distribution of node-level IVT over all sampled firing node–timesteps. Dashed lines indicate the AR category (Cat 1–5) thresholds based on magnitude alone; see Ralph et al. [25] .
Figure 3: Concept 99 dialed over a 5-day rollout of global ARs (initialized 13 Nov 2021, four years past the SAE training period). Panels show within-AR (A) IVT, (B) column water vapor, (C) 10m wind, and (D) precipitation.
Figure 4: All 4,096 concepts from each SAE plotted by breadth (firing rate) vs. correlation with AR intensity; Matryoshka SAE (top) and Standard SAE (bottom) at layers 0, 8, 15. Colors correspond to Matryoshka prefix groups. Each panel’s top AR concept and child concept Concept 3153 are labeled.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Standard Top- K
Matryoshka Top- K
Input dimension d
512
512
Dictionary size n
4096
4096
Active latents k
32 (per sample)
32 (per sample)
Auxiliary latents (AuxK)
512
512
Input normalization
mean-center, ℓ2 row
min–max [−1,1] , running-avg
Batch size
8192
4096
Appendix
Table 1: Training hyperparameters. All three depths of a given architecture use identical settings; only the input activation record differs. Both architectures use per-sample Top- K selection and an AuxK dead-latent revival term.
Figure 5: Training reconstruction loss for the six SAEs against fraction of training, one panel per architecture, colored by depth. Each loss is measured in that SAE’s own operating space, so the panels are not comparable in absolute value. The Matryoshka SAE logs every 500 steps, so its curve is drawn smoothed (bold) over the raw per-step values (faint) to show the trend.
Deep learning weather prediction models achieve remarkable predictive skill yet remain largely opaque: we know little about how they represent physical climate phenomena internally. Mechanistic interpretability through Sparse Autoencoders (SAEs) offers a principled route to decomposing these representations, but existing SAEs assume strictly linear feature superposition - a constraint ill-suited for the highly nonlinear atmospheric dynamics encoded in modern transformers. We introduce KAN-SAE, a sparse autoencoder whose encoder replaces the standard ReLU with learnable per-feature B-spline activations drawn from Kolmogorov-Arnold Networks (KANs), allowing each latent dimension to develop its own nonlinear gating profile. Applied to Sonny, KAN-SAE discovers 975 alive features (vs. 566 for a linear baseline, a 72% improvement) with 20% lower inter-feature redundancy and comparable reconstruction fidelity. Without any climate supervision, KAN-SAE identifies an interpretable European heatwave feature spatially concentrated over western Europe, and a western Pacific typhoon tracker confirmed by causal steering experiments. Our results demonstrate that nonlinear activations are essential for mechanistic interpretability of deep learning weather prediction models, recovering climate features that remain invisible to linear baselines.
Minjong Cheon
Department of Computer Science and Engineering Sejong University
Artificial Intelligence (AI) weather models are improving rapidly, and their forecasts are already competitive with long-established traditional Numerical Weather Prediction (NWP). To build confidence in this new methodology, it is critical that we understand how these predictions are generated. This is a huge challenge as these AI weather models remain largely black boxes. In other areas of Machine Learning (ML), mechanistic interpretability has emerged as a framework for understanding ML predictions by analysing the building blocks responsible for them. Here we present an open-source, highly adaptable tool which incorporates concepts from mechanistic interpretability. The tool organises internal latent representations from the model processor and allows for initial analyses, including cosine similarity and Principal Component Analysis (PCA), enabling the user to identify directions in latent space potentially associated with meteorological features. Applying our tool to the graph neural network GraphCast, we present preliminary case studies for mid-latitude synoptic-scale waves and specific humidity. These demonstrate the tool's ability to identify linear combinations of latent channels that appear to correspond to interpretable features.
Kirsten I. Tempest, Matthias Beylich, George C. Craig
Meteorological Institute Munich, Ludwig-Maximilians-Universität, 80333 Munich, Germany · Deutsches Zentrum für Luft- und Raumfahrt, Oberpfaffenhofen, Germany
ML foundation models are able to emulate atmospheric dynamics accurately and efficiently but operate as opaque ``black boxes''. We investigate the internal representations of the Aurora model using spatially pooled PCA and layer-wise relevance propagation (LRP). We find evidence that Aurora's latent space is primarily organized by seasonal cycles, whereas extreme storm events do not form a linearly separable cluster. LRP indicates that the model attends to features consistent with the 3D vertical structure of the Great Storm of 1987. Perturbation tests show masking relevant regions degrades forecasts 3.31× more than random masking. These findings suggest that Aurora learns meteorological coherence and vertical structure without explicit instruction.