We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This "recovery principle" yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different data share features prevalent in both, and sufficiently large SAEs recover both parent and child features. In contrast to conventional wisdom that SAE features are unstable and "split" as size increases, we find that these predictions hold on SAEs of sizes ranging from 512 to 131,072 trained on two large embedding models. From a theoretical perspective, our results suggest the promise of a scientific theory of representations based on atomic features. Practically, our results suggest the promise of scaling SAEs.
Figures & tables
Figure 1 : The recovery principle says that SAEs of increasing size recover a growing prefix of atomic features. This is consistent with observed variation in SAEs, while giving testable predictions that counter conventional wisdom: SAEs of different size and trained on different data share features.
Figure 2 : Testing P1, P2, and P3—the predictions of the atomic theory.
Figure 3 : Hierarchical feature recovery. (a) Top: parent and child F1 for a-prefixes. Parent recovery follows a U-shape. Bottom: recall of the selected a-prefix feature, broken down by prefix. (b) Extracted Food and Music hierarchies in the Gemini 131K SAE. Activating examples for each feature are given in Tables 8 and 9 .
Table 1 : A polysemantic “transform”-related feature in the 4K gemini SAE, represented as a molecule of 131K features. It activates on mentions of “transform” as well as strings like “are autobots stronger than decepticons” (referencing the Transformers franchise). Examining activating examples in the table, the 131K features (treated as proxy atoms) seem comparatively monosemantic (e.g., representing the Transformers franchise). More examples for each atom are given in Table 11 .
Table 2 : Empirical a-prefix feature splitting viewed in terms of molecules. A family molecule splits into parent-child molecules into atoms. In the middle regime, no molecule corresponds to a feature that activates on the full family (feature splitting). Such a feature exists in the small and large regimes.
Appendix figures & tables46 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : A sketch of how the tree construction completes the proof. Removing one vertex r from C , there exists a matching denoted by the solid lines from C∖{r} to N(C) (by the “Hall minimality” of C ). We can then construct a tree on C where the existence of an edge ij implies that m(j)∈Fi (so Fi∩Fj is non-empty). Splitting the tree into its two parts, this gives us a subset of C that has a large neighborhood, providing the desired contradiction.
Figure 5 : When dim(BM′,K′,S)<dim(AM), the set of (M,K) -factorizations that also have an (M′,K′) -factorization is measure zero under a generic distribution over (M,K) -factorizations.
Source collection
# embeddings
Emotion [ 64 ]
19,930
FEVER [ 65 , 66 ]
11,079,420
GooAQ [ 67 ]
6,024,992
HotpotQA [ 68 ]
10,564,510
MS MARCO [ 69 ]
9,351,785
Natural Questions [ 70 ]
200,462
Appendix
Table 3: SAE training datasets.
Width or subset
k
512
32
1,024
32
2,048
32
4,096
32
8,192
64
16,384
64
Appendix
Table 4: Sparsities for the 26 SAEs in the main width and distribution experiments. Subset runs all use m=16,384 and k=64 . The Gemini width-131,072 run uses 2M steps and auxiliary coefficient 1.0; every other listed run uses 1M steps and coefficient 0.25.
Control
Model
Width
k
Fixed k
Gemini
512
128
Fixed k
Gemini
4,096
128
Fixed k
Gemini
32,768
128
Fixed k
Nemotron
512
128
Fixed k
Nemotron
4,096
128
Fixed k
Nemotron
32,768
128
Appendix
Table 5: The nine additional SAEs used for the fixed-sparsity and alternate-seed controls. All use 1,000,000 optimizer steps, batch size 2048, and auxiliary coefficient 0.25.
Model
Original width/ k
Same-width cross-sparsity matches
Original source persistence
k=128 source persistence
Gemini
512/32
389 (76.0%)
303 (59.2%)
221 (43.2%)
4,096/32
2,949 (72.0%)
2,974 (72.6%)
2,405 (58.7%)
32,768/64
16,466 (50.3%)
15,881 (48.5%)
12,893 (39.3%)
Nemotron
512/32
410 (80.1%)
259 (50.6%)
210 (41.0%)
4,096/32
3,286 (80.2%)
2,841 (69.4%)
2,607 (63.6%)
32,768/64
18,198 (55.5%)
12,750 (38.9%)
13,609 (41.5%)
Appendix
Table 6: Fixed-sparsity source replacement, using the same full-sweep target sets as the seed control. Only the source changes from its original k to k=128 ; all original larger-width targets are retained. Same-width counts match the original and replacement source dictionaries. All comparisons use signed decoder cosine similarity ≥0.7 ; percentages divide by source width.
Width/ k
Same-width cross-seed matches
Original source persistence
Alternate source persistence
512/32
470 (91.8%)
303 (59.2%)
293 (57.2%)
4,096/32
3,564 (87.0%)
2,974 (72.6%)
2,979 (72.7%)
32,768/64
17,856 (54.5%)
15,881 (48.5%)
15,998 (48.8%)
Appendix
Table 7: Gemini direction stability across training seeds. Persistence uses the independent pairwise matching and intersection criterion.
Figure 6 : Persistent stability under strict exclusive matching at 0.7 . A source direction counts only if it has a signed-cosine match above 0.7 in every larger trained dictionary, with no competing partner above 0.7 at either endpoint in each comparison.
Figure 7 : Persistent stability across SAE widths using activation similarity at signed Pearson correlation at least 0.7 . Left: the number of source features matched in every independently computed pairwise maximum matching to a larger SAE on the retained candidate graphs. Right: this count divided by source width. Both panels compare Gemini and Nemotron, using the fixed tie-breaking procedure in Section C.2.3 .
Figure 8 : Threshold sensitivity of direction persistence.
Figure 9 : Threshold sensitivity of activation persistence.
Figure 10 : Prevalence across SAE widths (left) and KMeans cluster counts (right), for Gemini (top) and Nemotron (bottom). SAE prevalence uses activation strictly greater than each SAE’s pooled median positive activation; KMeans prevalence uses nearest-centroid assignment. KMeans models are trained on the full corpus for 100 iterations. Colors indicate dictionary size. Curves show feature-count-scaled densities of log10 prevalence, with shared axes and a square-root vertical scale. Zero-rate features are omitted from the logarithmic axis.
Figure 11 : Prevalence for SAEs with fixed TopK sparsity k=128 (left) and KMeans (right), for Gemini (top) and Nemotron (bottom). SAE widths are 512, 4,096, 32,768, 65,536, and 131,072; KMeans cluster counts are 512, 4,096, 16,384, and 131,072. SAE prevalence is the fraction of examples with activation strictly above each SAE’s pooled median positive activation; KMeans prevalence is nearest-centroid assignment frequency after 100 full-corpus training iterations. Colors indicate the number of features or clusters. Curves show feature-count-scaled densities of log10 prevalence with bandwidth 0.12 , shared axes, and a square-root vertical scale. Zero-rate features are omitted from the logarithmic axis.
Figure 12 : Median-threshold feature prevalence and direction persistence under independent pairwise matching and intersection in Gemini (left) and Nemotron (right). Each curve represents a source width in the original sparsity schedule. Points group source features into up to ten prevalence-quantile bins, keeping ties together, and show the median prevalence and fraction selected in the persistent matching within each bin. All source features are included. Bin sizes can differ; lines connect descriptive bin summaries. The horizontal axis is logarithmic above 10−8 with a linear segment retaining zero-prevalence bins.
Figure 13 : Cross-distribution SAE and PCA matches under the alternative exclusive criterion at t=0.7 , using signed cosine for SAEs and absolute cosine for PCA. Each counted pair exceeds the threshold and has no other above-threshold partner at either endpoint. Entries divide the pair count by the dictionary width. Compared with the maximum-cardinality matching in Figure 2(b) , this criterion discards all pairs involving competing above-threshold partners. It therefore gives a stricter measure of unambiguous agreement between dictionaries.
Figure 14 : Cross-distribution SAE activation matching for Gemini (left) and Nemotron (right). Each cell reports the maximum number of one-to-one matches at signed Pearson correlation ≥0.7 , divided by the SAE width of 16,384. Both SAEs are evaluated on the row’s distribution; columns identify the comparison SAE’s training distribution.
Figure 15 : Median-threshold prevalence in both distributions and cross-distribution direction matching for Gemini (left) and Nemotron (right). Curves identify the source and target training distributions. The horizontal coordinate is the bin median of the minimum prevalence of a source feature across the two distributions; the vertical coordinate is the fraction selected in a one-to-one direction matching. Quantile binning preserves ties and includes all source features, with the same axis convention as Figure 12 .
Figure 16 : Hierarchical feature recovery at fixed k=128 . Mean held-out test F1 for Prefix, Cities, and GBIF at five SAE widths, pooling Gemini and Nemotron. Each curve averages the same category–model pairs as the corresponding curve in Figure 2(c) ; each category’s feature and threshold are selected using training rows only.
Figure 17 : Recovery of individual categories across SAE widths. A category is recovered when its selection-chosen detector achieves held-out test F1 >0.6 . Left axes show proportions; right axes show raw counts among 242 word prefixes, 93 geographic categories, and 499 organism categories. Geography uses different continent and country prompts.
Figure 19 : Gemini F1 recovery for 25 letter prefixes. Each panel is the analog of the upper panel of Figure 3(a) for one prefix from a to z , excluding x , across the original nine SAE widths and sparsity schedule. Blue curves show parent F1; bold green curves show the unweighted mean F1 of eligible two-letter children; faint green curves show individual children.
Figure 20 : Gemini F1 recovery for 25 letter prefixes at fixed k=128 . The analog of Figure 19 at five SAE widths: 512, 4,096, 32,768, 65,536, and 131,072. Panels cover a – z , excluding x . Blue curves show parent F1; bold green curves show the unweighted mean F1 of eligible two-letter children; faint green curves show individual children. Each category’s detector is selected using selection rows and evaluated on held-out test rows, with no F1 filtering.
Figure 21 : Learned hierarchies for Music and Food. The labels here are automatically generated using Codex; activating examples for each feature in the hierarchy are given in Tables 8 and 9 .
Table 8 : Food hierarchy with five unique activating examples per feature.
Table 9 : Music hierarchy with five unique activating examples per feature.
Figure 22 : Width-normalized prefix recovery across frequency distributions. Generating sparsity K=8 and SAE sparsity k=8 are fixed; each curve corresponds to a different α . All models train on 1,024M examples. Rows show flat and hierarchical data; columns show geometric and activation recovery. The vertical axis is approximate prefix length in atoms divided by SAE width (three atoms per complete family for hierarchical data), with a 10% miss allowance. Lines show three-seed means and bands show seed ranges.
Figure 23 : Width-normalized prefix recovery across SAE sparsities. Generating sparsity K=8 and frequency exponent α=1.6 are fixed; curves vary SAE sparsity k . All models train on 1,024M examples. Panels, normalization, and the 10% prefix miss allowance follow Figure 22 . Lines show three-seed means and bands show seed ranges.
Figure 24 : Persistent-feature proportions across frequency distributions at t=0.8 . The intersection of source sets matched by independent maximum matchings to every larger trained width, divided by source SAE width. Generating sparsity K=8 and SAE sparsity k=8 are fixed; colors vary α . Rows show flat and hierarchical data; columns show signed decoder cosine and activation Pearson similarity. All models train on 1,024M examples. Lines show three-seed means and bands show seed ranges. The dashed α=0.8 curve is the finite-distribution control outside the theorem’s guarantees.
Figure 25 : Persistent-feature proportions across SAE sparsities at t=0.8 . The intersection of source sets matched by independent maximum matchings to every larger trained width, divided by source SAE width. Generating sparsity K=8 and frequency exponent α=1.6 are fixed; colors vary SAE sparsity k . Rows show flat and hierarchical data; columns show signed decoder cosine and activation Pearson similarity. All models train on 1,024M examples. Lines show three-seed means and bands show seed ranges.
dict. size
i
a^i with molecule illustration
z^i(cat)
z^i(dog)
A^z^(cat)
A^z^(dog)
1
1
62apet+61acat+61adogpetcatdog
63
63
22apet+acat+adog
2
1
21apet+21acatpetcat
2
0
apet+acat
apet+adog
2
21apet+21adogpetdog
0
2
3
1
apetpet
1
1
apet+acat
apet+adog
2
acatcat
1
0
3
adogdog
0
1
Appendix
Table 10 : Learned directions illustrate molecular splitting. Corresponding learned features illustrate mesoscale feature splitting and absorption: at size 2 only, there is no learned feature that is active on both cat and dog.
Table 11 : Activating examples for the five 131K atoms. Each column contains twelve examples, sorted by decreasing activation. These appear more monosemantic than the feature in Table 1 .
Figure 26 : Decoder direction similarities among a-prefix probes. Signed pairwise cosine similarities for Gemini SAEs of widths 32,768 ( k=64 , left) and 131,072 ( k=128 , right). At each width, we take the best selected feature for each of the 23 available one- and two-letter a-prefix categories, without an F1 filter, and deduplicate feature IDs. This yields 22 unique decoder directions: ac and ak share a probe at both widths. Rows and columns use the same prefix order in both panels, with a common color scale from −0.4 to 0.4 . Diagonal self-similarities equal 1 and are shown at the maximum red. The median cosine over the 231 unique off-diagonal pairs is 0.1438 at width 32,768 and 0.0084 at width 131,072.
Figure 27 : Manifolds arising from small sets of 131K gemini SAE activations. Top left shows PCA of raw embeddings, top right shows PCA of the SAE feature representation restricted to a subset of coordinates, and the bottom shows the activation pattern on these selected features. Importantly, the SAE-induced manifolds arise purely from correlation in activations, not correlation in feature directions.
Table 12 : Top-activating text and image examples for eight color-selective features in the Gemini 131K sparse autoencoder. Each row shows the five highest-activating distinct general-corpus texts and five highest-activating WikiArt images for one feature, ordered by decreasing activation within each corpus. Parentheses give raw post-TopK activations. Texts are truncated for display.
Figure 28 : Feature-activation matching between Gemini and Nemotron SAEs of equal width. Curves give maximum-cardinality one-to-one matching at signed Pearson thresholds t∈{0.5,0.6,0.7,0.8} on 89,227,558 shared examples. Left: matched-pair count. Right: count divided by the common SAE width, including dead and unmatched features in the denominator. Counts are exact on the retained candidate graphs and lower bounds on exhaustive matching. The 131K comparison is included because matching across models does not require a larger SAE.
Table 13 : Clarifying the meaning or scope of a term. Pearson r=0.352 .
Sparse Autoencoders (SAEs) have found success parsing neural representations into interpretable concepts, providing a basis for understanding and control. However, what exactly SAEs extract, and, correspondingly, the scientific conclusions we can draw from them, are not obvious. Empirically, the proof is in the pudding: SAEs learn interpretable features. Theoretically, we lack a clear account of what properties a 'concept' must satisfy for an SAE to extract it. There has been extensive identifiability work studying the conditions under which sparse coding recovers ground-truth features; however, these approaches tends to focus on simple data-generating models (e.g. sparse independent features) which poorly approximate the internet-swallowing language-model representations on which SAEs are trained. Here, avoiding data-generating models, we ask simply what properties any dictionary learning optimum must satisfy. Concretely, we extend local optimality analyses (Gribonval & Schnass, 2010) to the nonnegative joint-optimisation problem that vanilla SAEs approximate, and derive constraints relating optimal SAE features to their distributions. We use these constraints to explain a range of observed SAE behaviours - hierarchical splitting & absorption, the structure of residuals, and dense antipodal features - each reflecting how L1+nonnegativity interact with data to structure optimal dictionaries. Finally, we construct a novel large-dictionary convex problem and explore the wide atom-per-datapoint limit. In sum, we hope to tease model assumptions from unexpected observations, letting us learn more from SAEs' successes and provide principles for designing their successors.
Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts. A core SAE training hyperparameter is L0: how many SAE features should fire per token on average. Existing work compares SAE algorithms using sparsity-reconstruction tradeoff plots, implying L0 is a free parameter with no inherently correct value aside from its effect on reconstruction. In this work we study the effect of L0 on SAEs, and show that if L0 is not set correctly, the SAE fails to disentangle the underlying features of the LLM. If L0 is too low, the SAE will mix correlated features to improve reconstruction. If L0 is too high, the SAE finds degenerate solutions that also mix features. Further, we present a proxy metric that can help guide the search for the correct L0 for an SAE on a given training distribution. We show that our method finds the correct L0 in toy models and coincides with peak sparse probing performance in LLM SAEs. We find that most commonly used SAEs have an L0 that is too low. Our work shows that practitioners must set L0 correctly to train SAEs with monosemantic features.
Sparse autoencoders (SAEs) decompose language model activations into sparse features, yet these models traditionally encode each token independently, failing to expose information that persists across a sequence. We first show that temporal persistence can naturally emerge in standard SAE features: after a feature activates, the hidden state remains aligned with its direction, and past activations help reconstruct later hidden states. How long this lasts varies widely across features. We therefore introduce Persistent Sparse Autoencoders (Persistent SAEs), an extension of standard SAEs that learns a persistence coefficient for each feature, allowing the model to learn feature-specific timescales from reconstruction alone. Our experiments show that Persistent SAEs retain competitive reconstruction quality while learning a spectrum of timescales: short-timescale (fast) features stay locally interpretable, whereas long-timescale (slow) features accumulate information that identifies the current context. Moreover, we show in a prompt-injection monitoring case study that slow features preserve injection-related signals and remain causally effective over long contexts. These results suggest that Persistent SAEs offer new opportunities for interpreting and monitoring language models via persistent sparse features.