Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We introduce a family of chunk-level SAEs that encode mean-pooled activations over chunks, each a contiguous span of tokens: Mean-Chunk reconstructs the observed chunk, Cross-Chunk predicts an independently processed neighbor, and Joint-Chunk combines both targets. These designs separate the effect of a larger observation unit from that of predicting information shared across passages. With matched training data, chunk-level SAEs remain powerful interpretability tools while learning reliable semantic features that capture high-level concepts and respond selectively to relevant content. Their strengths are complementary: Mean-Chunk improves high-level feature discovery, reasoning detection beyond surface cues, and steering; Cross-Chunk leads document retrieval and classification transfer while producing selective, persistent features. Changing what an SAE sees and predicts yields reliable semantic features for more meaningful tasks. We demonstrate their practical value through gains across downstream tasks such as retrieval, reasoning detection, and steering.
Figures & tables
Figure 1: SAE representation quality and downstream utility. Left: reconstruction fidelity (activation preservation), semantic interpretability (concept clarity and abstraction), feature persistence (consistency across related contexts), and dictionary utilization (breadth of feature use). Right: retrieval (semantic matching), classification transfer (domain generalization and label efficiency), reasoning detection (sensitivity beyond surface cues), and steering (causal output control). Chunk-level SAEs learn more meaningful features with complementary strengths across downstream tasks.
Figure 2: Rethinking token-level sparse autoencoders. Top left: An arithmetic example illustrates we often get ’boring’ features. Top right: The same feature activates on “wait” in both backtracking and non-reasoning contexts. Bottom: We investigate whether modifying input granularity and training objectives yields more meaningful SAE features and validate across diverse downstream applications.
Figure 3: Training fidelity and feature interpretation. Top (a): RFVE during training, normalized to a target-appropriate reference. Bottom (b): architectures grouped by encoding granularity, focused interpretability (InterpScore), and high-level feature fraction (HighLevelFraction). Chunk-level SAEs improve semantic feature quality while retaining strong target-relative fidelity.
Figure 4: Changes and interpretation of top-5 feature activation values for each SAE across texts from different domains. Left: normalized traces of each SAE’s own top-five features across programming, medical, news, and mathematics passages. Right: automatic interpretations matched by color. Feature identities differ across SAEs. Cross-Chunk SAE exhibits sharper semantic selectivity, activating domain-relevant features while leaving unrelated features largely inactive.
Figure 5: Same-document retrieval beyond lexical overlap. Left (A): Recall@5 for word-set Jaccard matching and SAE representations; the dashed line marks the dense hidden-state baseline. Right (B): an example about volcanic activity with no shared content words, showing each SAE’s rank for the correct partner among length-matched candidates. Chunk-level SAEs outperform token-level baselines in retrieval, with Cross-Chunk SAE delivering substantial further gains.
Figure 6: Reasoning detection beyond surface cues. (A) Native reasoning recall on original texts with unrelated Pile controls; (B) recall with reasoning structure preserved but surface cues removed; (C) false feature activation with cues retained but reasoning structure removed. Mean-Chunk retains strong native recall, detects cue-free reasoning, and rejects cue-only distractors.
Figure 7: Classification transfer and feature steering. Left (A): low-label AUC across supervision budgets (solid bars) and OOD accuracy on future-year documents (hatched bars). Right (B): steering scores combining feature concept and coherence preservation; 50 denotes neutral output. Cross-Chunk leads both classification metrics, while Mean-Chunk achieves the highest steering score.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Value
Model layer / hidden width
21 / 4096
Dictionary width / BatchTopK budget
65,536 / 128
Training token occurrences
1,000,000,000
Validation / test occurrences
10,000,128 / 10,000,128
Training updates / global token batch
31,250 / 32,000
Chunk lengths
32, 64, 128, 256, 512
Appendix
Table 1: Shared model, training budget, and chunk-pair caches for all five SAE families and every Joint-Chunk partner weight. Token occurrences account for data volume; each objective retains its own encoding and loss-averaging unit for the token-level or chunk-level prediction task.
Figure 8: Training stability and sparse-code activity. (A) Training-batch and fixed-monitor RFVE, shown as logged values and nine-observation moving averages. (B) Validation effective L0 , the mean number of active coordinates per code. (the sparsity variation curve demonstrates stability.) (C) Fraction of dictionary coordinates that have not activated in the preceding ten million token occurrences. (D) Validation empty-code rate; the symmetric-log axis has a linear region below 0.001%. Coincident curves use staggered markers. For Joint-Chunk, the activity summaries show the smaller effective L0 and larger empty-code rate across its full-code and partner-prefix views.
Figure 9: Scientific-domain structure in document representations. Independently fitted t-SNE projections of the same 2,048 held-out ArXiv document codes for BatchTopK, Temporal, Mean-Chunk, and Cross-Chunk. Colors identify domains. Titles and callouts report purity in the cached 50-dimensional cosine neighborhoods, using non-self ranks 2–11. Cross-Chunk has the highest overall purity among the displayed methods, with domain-specific values highlighted in the callouts.
Figure 10: Semantic neighborhoods across domains and time. (A) Cosine-neighborhood label purity by ArXiv domain. (B) Filled circles show within-test purity using non-self ranks 2–11; open diamonds show OOD purity over the ten nearest test codes. Both query collections contain 2,048 documents balanced across eight classes. All statistics use label-free 50-D SVD codes.
Method
Silhouette
Neighbor purity
Cross-time NN
BatchTopK
0.086
0.667
0.745
Temporal
0.102
0.678
0.744
Mean-Chunk
0.068
0.666
0.736
Cross-Chunk
0.100
0.706
0.774
Appendix
Table 2: Separation, neighborhood agreement, and cross-time nearest-neighbor classification for frozen ArXiv codes. Higher is better; bold marks column maxima across the displayed methods.
Figure 11: Classification with limited supervision. (A) Test accuracy across seven labels-per-class budgets. (B) Within-seed differences from BatchTopK in percentage points. Points average five shared sampling seeds, and bands show one standard error of the mean, computed from paired differences in panel B. The curves expand the aggregate low-label summary in Section 5.3 .
Figure 12: Reasoning decisions under matched pooling and recalibration. Matched minus native-resolution rates for BatchTopK and Temporal: native recall, cue-free recall, and cue-only rejection. Comparisons pair the same 100 documents per relation and view; whiskers show paired-bootstrap 95% intervals from 20,000 resamples. Mean-Chunk and Cross-Chunk retain the same decisions on these views. Positive differences indicate higher reasoning recall or cue-only rejection.
α
Self FVE
Partner FVE
Composite RFVE
0.25
0.9263
0.6698
0.9355
0.50
0.9163
0.6701
0.9350
1.00
0.9045
0.6662
0.9367
1.50
0.8966
0.6621
0.9378
Appendix
Table 3: Target fidelity across Joint-Chunk weights. Each setting is evaluated on the shared monitor. Self and partner FVE describe the two prediction targets, and composite RFVE uses Eq. 24 .
Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts. A core SAE training hyperparameter is L0: how many SAE features should fire per token on average. Existing work compares SAE algorithms using sparsity-reconstruction tradeoff plots, implying L0 is a free parameter with no inherently correct value aside from its effect on reconstruction. In this work we study the effect of L0 on SAEs, and show that if L0 is not set correctly, the SAE fails to disentangle the underlying features of the LLM. If L0 is too low, the SAE will mix correlated features to improve reconstruction. If L0 is too high, the SAE finds degenerate solutions that also mix features. Further, we present a proxy metric that can help guide the search for the correct L0 for an SAE on a given training distribution. We show that our method finds the correct L0 in toy models and coincides with peak sparse probing performance in LLM SAEs. We find that most commonly used SAEs have an L0 that is too low. Our work shows that practitioners must set L0 correctly to train SAEs with monosemantic features.
Sparse autoencoders (SAEs) decompose language model activations into sparse features, yet these models traditionally encode each token independently, failing to expose information that persists across a sequence. We first show that temporal persistence can naturally emerge in standard SAE features: after a feature activates, the hidden state remains aligned with its direction, and past activations help reconstruct later hidden states. How long this lasts varies widely across features. We therefore introduce Persistent Sparse Autoencoders (Persistent SAEs), an extension of standard SAEs that learns a persistence coefficient for each feature, allowing the model to learn feature-specific timescales from reconstruction alone. Our experiments show that Persistent SAEs retain competitive reconstruction quality while learning a spectrum of timescales: short-timescale (fast) features stay locally interpretable, whereas long-timescale (slow) features accumulate information that identifies the current context. Moreover, we show in a prompt-injection monitoring case study that slow features preserve injection-related signals and remain causally effective over long contexts. These results suggest that Persistent SAEs offer new opportunities for interpreting and monitoring language models via persistent sparse features.
Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE families remains untested. Single-token features that activate on one vocabulary item provide the diagnostic case where ground truth permits direct comparison. We analyze 3.9M features across six models and three SAE families using zero-ablation at full layer depth. Single-token features cluster 4.7x tighter in decoder space and concentrate in early layers (Layer 0 in GPT2-Small; L0-L4 in Gemma). Ablating them yields Benjamini-Hochberg-significant logit reductions in 178 of 208 full-layer conditions, with depth controlling whether damage cascades downstream or shapes the output directly. Cross-family causal differences exceed within-family scale effects: on the same base model, GemmaScope and BatchTopK features remain causally anchored, while LlamaScope features are locally redundant. The target token's rank recovers to within 2x baseline 96-98% of the time after the same ablation, and a controlled activation-function comparison reverses sign within the same model, leaving training recipe as the residual candidate. Cross-family interpretability claims are therefore sensitive to training methodology, not just activation function or scale.