Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions, distinguishing constructive from destructive interference and frequent weak interactions from rare strong ones. Under local fixed-support assumptions, we characterize how architectural constraints shape interference through four mechanisms: feature orthogonalization, bias compensation, gain adaptation, and encoder-decoder separation. Experiments with sparse autoencoders show that constrained architectures selectively reduce overlap among co-active features, while bias, gain, and encoder freedom allow constructive cross-contributions to remain. Together, these results show that interference in learned representations depends not only on feature geometry, but also on how features are used and on the architecture that produces their codes.
Figures & tables
Figure 1: Effective interference is a joint property of concept geometry and code statistics. It factors as Iij=ρijπijMij , where ρij=di⊤dj captures the geometric overlap between dictionary elements di,dj , while πijMij=E[zizj] captures how their corresponding codes zi,zj are used across the data, through co-activation frequency πij and conditional magnitude Mij .
Figure 2: Geometric and effective interference have distinct training dynamics. A) Global geometric overlap rises, whereas B) effective interference falls from its peak. Consistent with Proposition 4.3 , this reflects selective orthogonalization of co-active pairs, reducing realized interactions without globally orthogonalizing the dictionary. C) Residual coupling declines but remains nonzero, so fixed-support code optimality ( Proposition 4.1 ) is approached only approximately (see Figure 7 for joint distribution of decoder orientation ρij and co-activation-weighted magnitude πijMij ).
Figure 3: Architectural freedom shifts co-active feature pairs toward constructive decoder interactions. Joint distributions of orientation ρij and co-activation-weighted magnitude πijMij . In the tied, normalized, bias-free model, pairs with large πijMij concentrate near ρij=0 , consistent with local orthogonalizing pressure. With bias, encoder-gain freedom, or an untied encoder, some of these pairs extend toward positive ρij . This pattern recurs across other SAEs ( Figure 8 ).
Figure 4
Figure 6: SAE reconstruction contains substantial signed effective interference. Constructive I+ , destructive I− , and net Inet contributions across architectural variants and SAE families. Sizable positive and negative terms can cancel, so small net interference does not imply weak interactions. Relative to the tied case, architectural freedom generally shifts the balance toward higher interference.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
MMCS
median
MMCS∣⋅∣
>0.9
>0.7
JumpReLU
0.711
0.785
0.712
24.8%
61.1%
TopK
0.548
0.533
0.552
8.9%
31.2%
BatchTopK
0.436
0.335
0.444
6.0%
18.5%
Matryoshka
0.425
0.332
0.433
3.1%
15.8%
Appendix
Table 1: Cross-seed dictionary agreement for the runs in Figure 2 , averaged over seed pairs (0,1) , (0,2) , and (1,2) .
TopK
BatchTopK
Matryoshka
JumpReLU
Constraint
FVU
L0
Dead
FVU
L0
Dead
FVU
L0
Dead
FVU
L0
Dead
(i) Tied
.118
40.0
0.00
.124
40.9
0.00
.126
40.7
0.00
.100
45.6
0.02
(ii) Tied + Bias
.086
40.0
0.00
.085
40.1
0.00
.091
40.1
0.00
.086
45.5
0.48
(iii) Tied + Gain
.091
40.0
0.00
.092
40.0
0.12
.096
40.2
0.01
.087
45.4
0.62
(iv) Tied + Gain + Bias
.085
40.0
0.01
.084
40.0
0.02
.090
40.1
0.01
.085
45.4
0.90
(v) Untied
.083
40.0
0.54
.084
40.8
5.41
.088
40.6
6.05
.083
45.0
17.62
Appendix
Table 2: Holdout evaluation after 200M training tokens. “Dead” is the percentage of features that never activate on the holdout set.
Figure 7: Transition of interference factors as training progress. This figure complements Figure 2 .
Figure 9: The untying mechanism to separate constructive decoder interactions from encoder cross-talk is consistent across SAEs, including TopK, BatchTopK, Matryoshka, and JumpReLU. This figure complements Figure 5 .
Figure 10: Interference across sparsity and dictionary width. TopK SAEs trained on Pythia-160M, varying constraints, support size k , and width p . Left: varying k Right: varying p .
Figure 11: Effective interference on MNIST. The same qualitative pattern as in language models appears: interference is lowest in the most constrained tied setting and increases as architectural degrees of freedom are added. This figure complements Figure 6 .
Figure 12: Decoder orientation and co-activation on MNIST. Relaxing the constraints broadens the distribution of decoder correlations and increases co-activation among non-orthogonal feature pairs. This figure complements Figure 3 .
Figure 13: Effective interference for independently trained SAEBench SAEs. Gemma-2-2B layer 12 (left) and Pythia-160M layer 8 (right), across the four architectures.
Figure 14: Decoder orientation and co-activation for independently trained SAEBench SAEs. Gemma-2-2B layer 12 (top) and Pythia-160M layer 8 (bottom).
Sparse Autoencoders (SAEs) have emerged as a powerful paradigm for disentangling feature superposition in transformer-based architectures, enabling precise control via activation steering. However, the theoretical foundations of compositional steering -- the simultaneous activation of distinct semantic latents -- remain under-explored. The prevailing Linear Representation Hypothesis often abstracts away non-linear interference effects that arise in overcomplete dictionaries. We present a geometric framework for analyzing the instability of feature unions. Modeling the activation space as a high-dimensional sparse cone manifold, we derive an asymptotic compositional-collapse threshold under a spherical dictionary model, characterized by the Gaussian mean width (statistical dimension) of the signal cone. We further show that, in the high-bias regime, ReLU rectification converts microscopic correlation-induced variance fluctuations into a systematic drift that accumulates under composition, yielding interference growth consistent with a ratchet effect. We validate the predicted scaling trends on structured semantic features extracted from CLEVR, where hierarchical correlations accelerate the transition relative to random baselines. Together, our results highlight geometric constraints on the scalability of union-based steering and motivate composition mechanisms that explicitly manage interference beyond naive linear superposition.
Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability bounds for fixed supports under subgaussian noise, proving the sufficient dimension scales linearly (d=Oε(klogm)) rather than prior worst-case quadratic limits. We then validate these bounds across system parameters through Gaussian-tail approximations. These results quantify the geometric constraints of the linear representation hypothesis, providing a framework for evaluating sparse autoencoders, compositional generalization, and neural interpretability.
Enrico Vompa
Applied Artificial Intelligence Group Tallinn University of Technology, Estonia
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. In practice, however, feature entanglement leads to interference such that localized interventions can have unintended downstream effects. Motivated by the \textit{Independent Causal Mechanisms} principle, we propose to constrain internal features to be almost orthogonal. We argue that this promotes modular representations amenable to causal intervention. We formalize this problem by characterizing the gap between an idealized isolated intervention and its realized effect on model outputs in terms of feature interference. We upper-bound the propagation of feature interference in terms of the self-coherence of the feature dictionary, and relate this discrepancy to an explicit orthogonality regularization on the dictionary itself. Empirically, we show that this regularization enables more isolated interventions on mathematical reasoning concepts while preserving model performance. Our code is available under \texttt{https://github.com/mrtzmllr/sae-icm}.
Moritz Miller, Florent Draye, Bernhard Schölkopf
Max Planck Institute for Intelligent Systems, Tübingen, Germany · ETH Zurich, Switzerland · Max Planck ETH Center for Learning Systems +1