Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
Figures & tables
Figure 1: Overview of the workflow. We test PoS recoverability from token-level SAE activations, identify and compact PoS -relevant latent groups, and validate them on held-out and controlled data.
Figure 2: One-vs-rest probing performance for each PoS . F1 scores are computed with 5-fold cross-validation on the GUM training split.
Figure 3: Non-zero coefficients for each one-vs-rest classifier; results are color coded by PoS class (Open, Closed, Other).
Figure 4: Number of salient latents required to reach 95% coverage for each PoS category. For each tag c , kc95 denotes the smallest number of highest-coefficient latents needed to activate on at least 95% of gold tokens of that category. Lower values indicate compact groups.
Figure 5: Row-normalized confusion matrix of the multi-class PoS classifier trained on the compact feature set L⋆ . Each cell reports the percentage of tokens of a gold PoS category predicted as each class.
Template
P
R
F1
I see/saw [DET] [NOUN]
0.207
0.984
0.341
I see/saw [DET] [ADJ] [NOUN]
0.215
0.967
0.351
I see/saw [DET] [NOUN] [PUNCT]
0.224
0.971
0.365
I have/had ([DET]) [NOUN]
0.186
0.977
0.313
I have/had ([DET]) [ADJ] [NOUN]
0.200
0.973
0.332
I have/had ([DET]) [ADJ] [NOUN] [PUNCT]
0.212
0.976
0.348
Table 1: Pseudo-multilabel classification results on the controlled dataset.
Figure 6: Co-activations of selected latent groups on the held-out treebank test set. Diagonal values correspond to recall for each PoS ; off-diagonal values correspond to false positive rates on other PoS categories.
Figure 7: Active L ⋆ latents per relevant PoS in each template variant. Each square is one latent that is active on at least 25% of template’s examples. Color saturation indicates activation percentage (darker = more active).
Probe
Accuracy
Macro F1
SAE
0.88
0.78
Layer 30
0.92
0.84
Layer 0 (Embed.)
0.88
0.78
Table 2: Performances using different probes.
Overlap (%)
Evaluation
Accuracy
Macro-F1
0
Cross-validation
0.24
0.16
Train/test
0.23
0.16
25
Cross-validation
0.49
0.41
Train/test
0.49
0.41
50
Cross-validation
0.69
0.60
Train/test
0.69
0.60
Table 3: Accuracy and Macro-F1 across overlap levels, comparing cross-validation and train/test evaluation.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Per- PoS one-vs-rest classifier performances across layers.
Figure 9: Per- PoS one-vs-rest classifier top coefficients. Open class PoS es.
Figure 10: Per- PoS one-vs-rest classifier top coefficients. Closed class PoS es.
Figure 11: Per- PoS one-vs-rest classifier top coefficients. Other class PoS es.
POS
POS Type
Non-zero Latents
Frac. of Latent Space
NOUN
open
10043
0.0766
VERB
open
7404
0.0565
ADJ
open
8877
0.0677
PROPN
open
5252
0.0401
ADV
open
5218
0.0398
INTJ
open
1010
0.0077
Appendix
Table 4: Non-zero latents and fraction of latent space by POS tag.
PoS Class
mean ± std.
Closed
3281.74 ± 1895.65
Open
7926.52 ± 2153.69
Other
2094.42 ± 221.91
Appendix
Table 5: Number of latents with non-zero coefficients for classification for each one-vs-rest classifier, aggregated by PoS Group. Classifiers for Open -class PoS tags have the highest number of non-zero coefficients.
Class
POS
Coverage over Salience (%)
Open
adj
0.91
adv
1.38
intj
3.76
noun
0.34
propn
1.14
verb
0.70
Appendix
Table 6: Coverage over number of salient features per UPOS, grouped by class.
precision
recall
f1-score
support
ADJ
0.85
0.84
0.85
11489
ADP
0.93
0.88
0.91
16655
ADV
0.84
0.79
0.81
8556
AUX
0.93
0.91
0.92
9682
CCONJ
0.96
0.97
0.96
5854
DET
0.95
0.93
0.94
14307
Appendix
Table 7: Multiclass classifier performances.
Figure 12: Coefficients importance heatmap for each PoS in the mutliclass classifier trained with 5-fold cross-validation on the GUM Training set.
Figure 13: Parameter sweep for C and τ for the compact feature classification task.
Figure 14: Density plot: number of activations from L(kc′⋆) over PoS with tag c′ in the Treebank test set, for all c′∈C .
Class
POS
Zero-act. %
Mcc
D(c)
Open
adj
0.056
0.944
0.305
adv
0.040
0.960
0.217
intj
0.076
0.924
0.093
noun
0.051
0.949
0.336
propn
0.027
0.973
0.273
verb
0.048
0.952
0.303
Appendix
Table 8: Zero-activation fraction, Mcc , and distinctiveness D(c) per UPOS category.
Figure 15: KDE of active-latent percentages per PoS tag in the I + [VERB] templates.
Figure 16: KDE of active-latent percentages per PoS tag in the I have/had templates.
Figure 17: KDE of active-latent percentages per PoS tag in the I + [VERB] templates.
Figure 18: KDE of active-latent percentages per PoS tag in the There is/are templates.
Figure 19: Demonstration of the behavior on first-token activations. We show that several activations shift from PRON in the top sentence (the first token “I”) to the first token NUM in the bottom sentence.