Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
Figures & tables
Figure 1: Overview of the workflow. We test PoS recoverability from token-level SAE activations, identify and compact PoS -relevant latent groups, and validate them on held-out and controlled data.
Figure 2: One-vs-rest probing performance for each PoS . F1 scores are computed with 5-fold cross-validation on the GUM training split.
Figure 3: Non-zero coefficients for each one-vs-rest classifier; results are color coded by PoS class (Open, Closed, Other).
Figure 4: Number of salient latents required to reach 95% coverage for each PoS category. For each tag c , kc95 denotes the smallest number of highest-coefficient latents needed to activate on at least 95% of gold tokens of that category. Lower values indicate compact groups.
Figure 5: Row-normalized confusion matrix of the multi-class PoS classifier trained on the compact feature set L⋆ . Each cell reports the percentage of tokens of a gold PoS category predicted as each class.
Template
P
R
F1
I see/saw [DET] [NOUN]
0.207
0.984
0.341
I see/saw [DET] [ADJ] [NOUN]
0.215
0.967
0.351
I see/saw [DET] [NOUN] [PUNCT]
0.224
0.971
0.365
I have/had ([DET]) [NOUN]
0.186
0.977
0.313
I have/had ([DET]) [ADJ] [NOUN]
0.200
0.973
0.332
I have/had ([DET]) [ADJ] [NOUN] [PUNCT]
0.212
0.976
0.348
Table 1: Pseudo-multilabel classification results on the controlled dataset.
Figure 6: Co-activations of selected latent groups on the held-out treebank test set. Diagonal values correspond to recall for each PoS ; off-diagonal values correspond to false positive rates on other PoS categories.
Figure 7: Active L ⋆ latents per relevant PoS in each template variant. Each square is one latent that is active on at least 25% of template’s examples. Color saturation indicates activation percentage (darker = more active).
Probe
Accuracy
Macro F1
SAE
0.88
0.78
Layer 30
0.92
0.84
Layer 0 (Embed.)
0.88
0.78
Table 2: Performances using different probes.
Overlap (%)
Evaluation
Accuracy
Macro-F1
0
Cross-validation
0.24
0.16
Train/test
0.23
0.16
25
Cross-validation
0.49
0.41
Train/test
0.49
0.41
50
Cross-validation
0.69
0.60
Train/test
0.69
0.60
Table 3: Accuracy and Macro-F1 across overlap levels, comparing cross-validation and train/test evaluation.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Per- PoS one-vs-rest classifier performances across layers.
Figure 9: Per- PoS one-vs-rest classifier top coefficients. Open class PoS es.
Figure 10: Per- PoS one-vs-rest classifier top coefficients. Closed class PoS es.
Figure 11: Per- PoS one-vs-rest classifier top coefficients. Other class PoS es.
POS
POS Type
Non-zero Latents
Frac. of Latent Space
NOUN
open
10043
0.0766
VERB
open
7404
0.0565
ADJ
open
8877
0.0677
PROPN
open
5252
0.0401
ADV
open
5218
0.0398
INTJ
open
1010
0.0077
Appendix
Table 4: Non-zero latents and fraction of latent space by POS tag.
PoS Class
mean ± std.
Closed
3281.74 ± 1895.65
Open
7926.52 ± 2153.69
Other
2094.42 ± 221.91
Appendix
Table 5: Number of latents with non-zero coefficients for classification for each one-vs-rest classifier, aggregated by PoS Group. Classifiers for Open -class PoS tags have the highest number of non-zero coefficients.
Class
POS
Coverage over Salience (%)
Open
adj
0.91
adv
1.38
intj
3.76
noun
0.34
propn
1.14
verb
0.70
Appendix
Table 6: Coverage over number of salient features per UPOS, grouped by class.
precision
recall
f1-score
support
ADJ
0.85
0.84
0.85
11489
ADP
0.93
0.88
0.91
16655
ADV
0.84
0.79
0.81
8556
AUX
0.93
0.91
0.92
9682
CCONJ
0.96
0.97
0.96
5854
DET
0.95
0.93
0.94
14307
Appendix
Table 7: Multiclass classifier performances.
Figure 12: Coefficients importance heatmap for each PoS in the mutliclass classifier trained with 5-fold cross-validation on the GUM Training set.
Figure 13: Parameter sweep for C and τ for the compact feature classification task.
Figure 14: Density plot: number of activations from L(kc′⋆) over PoS with tag c′ in the Treebank test set, for all c′∈C .
Class
POS
Zero-act. %
Mcc
D(c)
Open
adj
0.056
0.944
0.305
adv
0.040
0.960
0.217
intj
0.076
0.924
0.093
noun
0.051
0.949
0.336
propn
0.027
0.973
0.273
verb
0.048
0.952
0.303
Appendix
Table 8: Zero-activation fraction, Mcc , and distinctiveness D(c) per UPOS category.
Figure 15: KDE of active-latent percentages per PoS tag in the I + [VERB] templates.
Figure 16: KDE of active-latent percentages per PoS tag in the I have/had templates.
Figure 17: KDE of active-latent percentages per PoS tag in the I + [VERB] templates.
Figure 18: KDE of active-latent percentages per PoS tag in the There is/are templates.
Figure 19: Demonstration of the behavior on first-token activations. We show that several activations shift from PRON in the top sentence (the first token “I”) to the first token NUM in the bottom sentence.
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure. To probe this gap further, we study active latent sets under well-controlled semantic modifications, revealing a substantial mismatch between human judgements of conceptual change and change in the SAE active set. We interpret this as evidence that, outside idealised settings, SAE features do not compose via simple bag-of-features semantics.
Sparse autoencoders (SAEs) decompose language model activations into sparse features, yet these models traditionally encode each token independently, failing to expose information that persists across a sequence. We first show that temporal persistence can naturally emerge in standard SAE features: after a feature activates, the hidden state remains aligned with its direction, and past activations help reconstruct later hidden states. How long this lasts varies widely across features. We therefore introduce Persistent Sparse Autoencoders (Persistent SAEs), an extension of standard SAEs that learns a persistence coefficient for each feature, allowing the model to learn feature-specific timescales from reconstruction alone. Our experiments show that Persistent SAEs retain competitive reconstruction quality while learning a spectrum of timescales: short-timescale (fast) features stay locally interpretable, whereas long-timescale (slow) features accumulate information that identifies the current context. Moreover, we show in a prompt-injection monitoring case study that slow features preserve injection-related signals and remain causally effective over long contexts. These results suggest that Persistent SAEs offer new opportunities for interpreting and monitoring language models via persistent sparse features.
Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts. A core SAE training hyperparameter is L0: how many SAE features should fire per token on average. Existing work compares SAE algorithms using sparsity-reconstruction tradeoff plots, implying L0 is a free parameter with no inherently correct value aside from its effect on reconstruction. In this work we study the effect of L0 on SAEs, and show that if L0 is not set correctly, the SAE fails to disentangle the underlying features of the LLM. If L0 is too low, the SAE will mix correlated features to improve reconstruction. If L0 is too high, the SAE finds degenerate solutions that also mix features. Further, we present a proxy metric that can help guide the search for the correct L0 for an SAE on a given training distribution. We show that our method finds the correct L0 in toy models and coincides with peak sparse probing performance in LLM SAEs. We find that most commonly used SAEs have an L0 that is too low. Our work shows that practitioners must set L0 correctly to train SAEs with monosemantic features.