Interpreting language models remains challenging due to the existence of residual stream, which linearly mixes and duplicates features across adjacent layers, causing single-layer analyses to miss this cross-layer structure. Cross-layer sparse autoencoders (SAEs) address layer mixing but operate in continuous space, where concepts split across many neurons without clear boundaries. We introduce Cross-Layer Vector Quantized-Variational Autoencoder (CLVQ-VAE), a novel framework which maps representations from a lower layer to a higher layer through a discrete vector-quantization bottleneck, collapsing duplicated residual-stream features into compact, interpretable concept vectors. Our approach combines top-k temperature-based sampling with exponential moving average (EMA) codebook updates, providing controlled exploration of the discrete latent space while maintaining codebook diversity. Across both encoder- and decoder-based models on ERASER-Movie, Jigsaw, and AGNews, CLVQ-VAE outperforms clustering, single-layer vector quantized-variational autoencoder (VQ-VAE), and sparse autoencoder (SAE) baselines across three evaluation axes: removing identified concepts drops downstream probe accuracy by up to 93%, LLM judges rank our concepts first in 66.7% of comparisons, and human annotators recover model predictions from our visualizations with 78% accuracy versus 54% for clustering.
Figures & tables
Figure 1: Overview of the CLVQ-VAE framework for cross-layer concept discovery. Lower-layer activations are passed through an adaptive residual encoder, discretized via vector quantization into concept vectors, and decoded to reconstruct higher-layer representations.
Dataset
Label Purity
Label-Div. Tok.
Mean JSD
SL/DL Ratio
ERASER-Movie
0.691 (0.50)
75.8%
0.848
4.26×
Jigsaw
0.646 (0.50)
32.1%
0.916
7.56×
AGNews
0.353 (0.25)
43.7%
0.823
2.10×
Table 2: Codebook concept specificity (RoBERTa), with random-assignment baseline in parentheses. SL/DL Ratio denotes same- vs. different-label Jaccard overlap at TF-IDF cosine ≥0.1 .
Method
Mean Rating ± Std
MRR
Win Rate
CLVQ-VAE
1.890 ± 0.877
0.611
66.7%
Single-Layer
1.823 ± 0.874
0.458
44.4%
Cross-Layer SAE
1.800 ± 0.888
0.542
44.4%
Clustering
1.675 ± 0.855
0.472
44.4%
Table 3: LLM-judge based evaluation of methods across models and datasets.
Method
Fleiss’ Kappa ( κ )
Avg. Confidence
Model Alignment Rate
Clustering
0.59
5.981
54.14%
CLVQ-VAE
0.864
8.44
78.20%
Table 4: Human evaluation results comparing CLVQ-VAE with baseline clustering approach for ERASER-Movie dataset. Higher values indicate better performance across all metrics.
Method
Mean Rating ± Std
MRR
Win Rate
Spherical
1.903 ± 0.306
0.694
62.5%
K-means
1.841 ± 0.330
0.667
58.3%
Random
1.800 ± 0.390
0.472
25.0%
Table 6: LLM-judge based evaluation of initialization methods across models and datasets.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Train
Dev
Tags
ERASER
13878
856
2
JIGSAW
9000
800
2
AGNEWS
16000
1200
4
Appendix
Table 8: The data size of each benchmark used in the evaluation: the ERASER Sentiment dataset, Jigsaw Toxicity dataset, and the AGNEWS dataset.
Category
Component
Value
Architecture
Codebook size
400
Commitment cost ( β )
0.1
Decoder layers
6
Decoder attention heads
8
Feedforward dimension
2048
Dropout
0.1
Appendix
Table 9: Hyperparameters used across all experiments.
Dataset
Model
k=1
k=5
k=10
ERASER- Movie
RoBERTa
0.5389 ± 0.1052
0.4085 ± 0.0656
0.4361 ± 0.1664
BERT
0.6877 ± 0.0378
0.5024 ± 0.0739
0.4657 ± 0.0733
LLaMA
0.9084 ± 0.0012
0.9092 ± 0.0014
0.9092 ± 0.0007
Qwen
0.8828 ± 0.0018
0.8782 ± 0.0037
0.8731 ± 0.0014
Jigsaw
RoBERTa
0.9147 ± 0.0026
0.9164 ± 0.0027
0.9172 ± 0.0022
BERT
0.8581 ± 0.0167
0.8432 ± 0.0102
0.8335 ± 0.0064
Appendix
Table 10: SAE faithfulness (perturbed accuracy) across subspace dimension k . Lower is better (larger concept removal). Averaged over 3 seeds.
Model
Dataset
Layer Pair
Original CLS
Random Perturbed CLS
RoBERTa
ERASER-Movie
8–12
0.8777
0.8190
RoBERTa
Jigsaw
8–12
0.9121
0.9121
RoBERTa
AGNews
8–12
0.7275
0.6875
BERT
ERASER-Movie
8–12
0.8248
0.8237
BERT
Jigsaw
8–12
0.8995
0.8995
BERT
AGNews
8–12
0.7458
0.7433
Appendix
Table 11: Reference baseline values for faithfulness evaluation across different model-dataset combinations and layer pairs.
Table 13: Ablating the identified salient concept versus other active codebook vectors. Lower perturbed accuracy indicates greater importance. Random perturbed is an unstructured baseline.
Dataset
Kendall’s W
Agreement Level
Jigsaw
0.900
Strong
ERASER-Movie
0.475
Weak
AGNews
0.225
Weak
Overall Average
0.533
Moderate
Appendix
Table 14: Inter-judge agreement on baseline rankings measured by Kendall’s coefficient of concordance.
Dataset
Kendall’s W
Agreement Level
Jigsaw
0.910
Strong
ERASER-Movie
0.639
Moderate
AGNews
0.843
Strong
Overall Average
0.793
Strong
Appendix
Table 15: Inter-judge agreement on initialization method rankings measured by Kendall’s coefficient of concordance.
Dataset
Token
Example sentence
Vector
Majority (purity)
ERASER-Movie
entertainment
“the movie does not serve as a serious thriller nor as comic entertainment (because of its serious tone).” (neg.)
#137
Neg. 91%
“Tarantino twists this age-old genre to produce over-the-top entertainment.” (pos.)
#73
Pos. 97%
simply
“Matthew Modine is quite simply terrible.” (neg.)
#49
Neg. 93%
“Princess Caraboo is simply an eminently enjoyable entertainment.” (pos.)
#246
Pos. 94%
Jigsaw
thank
“We can take our time considering the wider issues. Thank you!” (non-tox.)
#265
Non-tox. 100%
“Thank you for blocking him. That guy has been vandalizing the page for at least 2 weeks.” (toxic)
#399
Toxic 62%
Appendix
Table 16: Top label-divergent tokens for ERASER-Movie, Jigsaw, and AGNews (RoBERTa). Each row shows one example sentence per label, the dominant codebook vector the token routes to, and the majority class among all tokens assigned to that vector (purity %).
SL/DL Ratio
Dataset
Model
Label Purity
Rand.
Label-Div. Tok.
Mean JSD
τ=0.1
τ=0.3
ERASER-Movie
RoBERTa
0.691
0.50
75.8%
0.848
4.26×
8.93×
BERT
0.667
31.5%
0.560
1.29×
3.89×
Qwen
0.694
44.0%
0.690
1.09×
2.11×
LLaMA
0.761
37.9%
0.628
1.05×
1.43×
Jigsaw
RoBERTa
0.646
0.50
32.1%
0.916
7.56×
30.56×
Appendix
Table 17: Full codebook concept specificity results across all twelve model-dataset configurations. Label-Div. Tok. : percentage of label-divergent content tokens. Mean JSD : computed over those tokens. SL/DL Ratio : same-label to different-label Jaccard overlap ratio at two TF-IDF cosine thresholds.
Figure 2: Label purity distributions for ERASER-Movie/RoBERTa (left) and Jigsaw/RoBERTa (right). The red dashed line marks the random baseline; the blue dashed line marks the VQ-VAE mean. The spike at purity =1.0 represents vectors firing exclusively on one class.
Figure 3: Word similarity (TF-IDF cosine) vs. codebook assignment overlap (Jaccard) for same-label (blue) and different-label (red) sentence pairs, for ERASER-Movie/RoBERTa (left) and Jigsaw/RoBERTa (right). Dashed trend lines are fitted separately per group. The growing divergence between same- and different-label trend lines confirms that the codebook discriminates by meaning above and beyond surface vocabulary.
Model
Dataset
Faithfulness
Cosine Sim.
(Full vs. Quantized-Only)
(Full vs. Quantized-Only)
RoBERTa
ERASER-Movie
0.0594 vs 0.0560
0.751 vs 0.924
RoBERTa
Jigsaw
0.6127 vs 0.5152
0.575 vs 0.484
RoBERTa
AGNews
0.0992 vs 0.1067
0.906 vs 0.976
BERT
ERASER-Movie
0.5311 vs 0.7560
0.479 vs 0.760
BERT
Jigsaw
0.7372 vs 0.8752
0.312 vs 0.666
Appendix
Table 18: Impact of cross-attention on faithfulness and codebook quality. Full denotes the standard CLVQ-VAE model, while Quantized-Only removes cross-attention so that the decoder receives only the quantized code sequence. Lower values indicate better performance for both metrics. Bold indicates better performance.
Alpha Strategy
Initial
Epoch 10
Epoch 30
Final
Best Val Loss
Adaptive (Limited)
1.97
198.5
216
210.6
0.033
Adaptive (Complete)
1.86
25.54
50.70
63.21
0.033
Fixed α =0.0
238.1
237.3
239.8
237.2
0.045
Fixed α =0.1
180.9
232.3
238.1
230.8
0.040
Fixed α =0.4
1.95
126.9
160.2
157.2
0.036
Fixed α =0.75
1.077
1.883
27.517
40.106
0.033
Appendix
Table 19: Alpha parameter analysis showing perplexity evolution across training epochs and final validation loss.
Commitment Cost ( β )
ERASER Perplexity
Jigsaw Perplexity
0.0
213.45
163.45
0.1
210.26
164.07
0.3
189.74
145.65
0.6
170.76
81.38
1.0
23.94
30.71
Appendix
Table 20: Impact of commitment cost β on validation perplexity across datasets.
Temperature
Top-k
Validation Perplexity
0.5
5
207.14
1.0
5
210.63
2.0
5
217.14
3.0
5
220.09
1.0
1
207.07
1.0
10
210.37
Appendix
Table 21: Impact of temperature and top-k parameters on codebook utilization (validation perplexity) for ERASER-Movie on RoBERTa model. Higher perplexity indicates more diverse codebook usage.
Top-k
Temperature
Perturbed CLS Accuracy
1
1.0
0.0911
10
1.0
0.0864
100
1.0
0.0817
400
1.0
0.0877
400
0.1
0.0806
400
1.0
0.0877
Appendix
Table 22: Impact of temperature and top-k values on concept identification performance for ERASER-Movie on RoBERTa model. Despite significant differences in sampling parameters, perturbed CLS accuracies remain within a narrow range (0.0783–0.0911).
Model-Dataset
Layer Pair
Perturbed CLS
Original CLS
Random Perturbed
RoBERTa–ERASER
0–4
0.5140
0.4988
0.5012
RoBERTa–ERASER
4–8
0.7069
0.5374
0.5269
RoBERTa–ERASER
8–12
0.0583
0.8777
0.8190
RoBERTa–Jigsaw
0–4
0.4962
0.4962
0.4962
RoBERTa–Jigsaw
4–8
0.1734
0.7692
0.7653
RoBERTa–Jigsaw
8–12
0.5853
0.9121
0.9121
Appendix
Table 23: Layer pair analysis showing that layers 8–12 capture the most meaningful transformations across all datasets.
Figure 4: False negative example (Model: 0, Ground Truth: 1) showing concept clusters related to imitation and parody.
Figure 5: False positive example (Model: 1, Ground Truth: 0) showing terms of moderate approval despite the overall negative sentiment.
Figure 6: True negative example (Model: 0, Ground Truth: 0) showing terms expressing absence or deficiency.
Figure 7: True positive example (Model: 1, Ground Truth: 1) showing positive descriptors for “the acting is superb from everyone involved”.