Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural networks. Recent work showed that these interventions routinely push representations off the model's natural distribution, and that such divergence is sometimes harmless and sometimes pernicious: it can recruit pathways the model never uses on natural inputs, so that an intervention produces the expected answer through the wrong mechanism. No method currently tells the two cases apart. We make this question testable by planting hidden pathways inside pretrained language models; the pathways are silent on every benchmark prompt by construction, so which interventions depend on them is known exactly. Across 72 configurations and 100,800 interventions on GPT-2 small, we find three things. (i) Nearest-neighbour and local-PCA distances at the intervention site, as used in prior work, score below chance (AUROC 0.35-0.47) at picking out interventions that give the right answer through a planted pathway. (ii) Hidden-Pathway Contribution (HPC), a label-free test that clamps downstream units to the regime of natural runs with the same output and measures how much of the decision disappears, flags pathway-dominated interventions with AUROC >= 0.99 when the pathway shows up as unit-level out-of-regime activity, but fails when every unit stays within its natural range, which we identify as the open problem. (iii) Optimised interventions actively seek hidden pathways: on a gender task, DAS routes 90-95% of its successes through planted pathways for three of four families, and a downstream on-manifold penalty cuts this share to under 5% at a cost of 6-11 points of success rate. In unmodified GPT-2, successful interventions show almost no unit-level out-of-regime reliance.
Figures & tables
Figure 1: Setting. An intervention replaces the site representation h by h^ . It may reach the right answer through the natural computation or through a pathway that natural inputs never activate. Site-level divergence looks only at h^ . HPC tests downstream whether the decision depends on activity outside the natural regime. In our benchmark the hidden pathway is planted, so the answer is known.
Detector
AUROC (all)
AUROC (successful)
FPR null-space
site kNN
0.651 ± 0.034
0.586 ± 0.034
94.8
site Mahalanobis
0.803 ± 0.038
0.751 ± 0.041
99.6
site local-PCA †
0.681 ± 0.030
0.618 ± 0.031
93.2
Algorithm 1 †
0.837 ± 0.016
0.868 ± 0.012
4.0
HPC (ours)
0.931 ± 0.007
0.937 ± 0.005
0.0
Table 1: Synthetic setting (5 seeds, 3,000 interventions each). AUROC for detecting trap-dependent interventions, over all and over successful interventions, and false-positive rate (%) on harmless null-space divergence at the threshold that catches 90% of pernicious cases. † Grant et al. [8] .
site-level
downstream
downstream + causal (ours)
Model
Task
Pathway
#cfg
#pos
kNN
Maha.
LPCA †
Alg. 1 †
Down-OM
HPC-S
HPC-L
HPC-Loc
GPT-2
Gender
silent
5
3203
0.74
0.80
0.75
0.86
0.92
1.00
1.00
0.99
GPT-2
Gender
hijack
5
3620
0.75
0.78
0.75
0.83
0.92
0.68
1.00
0.99
GPT-2
Gender
distrib.
5
3144
0.75
0.80
0.75
0.86
0.91
0.69
1.00
0.99
GPT-2
Gender
in-range
5
223
0.66
0.78
0.67
0.66
0.74
0.57
0.59
0.74
GPT-2
SVA
silent
5
410
0.84
0.91
0.84
0.94
0.99
1.00
1.00
0.99
Table 2: Detecting pathway-dominated interventions (planted pathway supplies > 50% of the decision margin), all interventions. Mean AUROC over configurations (sites and pathway layers; seed 0) with ≥ 5 positives; best per row in bold. † Measures used by Grant et al. [8] .
site-level
downstream
downstream + causal (ours)
Model
Task
Pathway
#cfg
#pos
kNN
Maha.
LPCA †
Alg. 1 †
Down-OM
HPC-S
HPC-L
HPC-Loc
GPT-2
Gender
silent
5
1099
0.38
0.69
0.38
0.65
0.69
0.83
0.83
0.82
GPT-2
Gender
hijack
5
1393
0.40
0.65
0.40
0.57
0.67
0.57
0.87
0.85
GPT-2
Gender
distrib.
5
832
0.35
0.73
0.35
0.69
0.68
0.62
0.87
0.86
GPT-2
Gender
in-range
3
71
0.46
0.60
0.46
0.31
0.48
0.55
0.54
0.57
GPT-2
SVA
silent
2
24
0.45
0.82
0.45
0.87
0.94
1.00
1.00
0.99
Table 3: Right answer, wrong mechanism. AUROC restricted to successful interventions (the counterfactual answer was produced); positives are those whose decision flips when the planted pathways are removed. Nearest-neighbour and local-PCA site distances fall below chance .
Figure 2: AUROC for wrong-mechanism successes on GPT-2, averaged over the two tasks, by pathway family. Site-level LPCA is below chance throughout. HPC-L is the best plotted detector in every family; for in-range its average hides chance-level performance on the gender task (Table 3 ).
main runs
mitigation runs
G: DAS
H: restricted
G ′ : DAS
J: DAS+pen.
H ′ : restr.
I: restr.+pen.
Task
Pathway
succ.
path.
succ.
path.
succ.
path.
succ.
path.
succ.
path.
succ.
path.
Gender
silent
91.2
90.2
13.5
100.0
98.0
98.5
87.0
4.5
52.5
100.0
41.0
100.0
Gender
hijack
90.7
90.0
26.1
100.0
94.0
100.0
88.0
4.7
66.0
100.0
17.0
100.0
Gender
distrib.
87.3
95.1
12.4
100.0
96.0
100.0
89.0
4.5
34.5
100.0
30.0
92.3
Gender
in-range
68.1
16.5
0.9
75.0
90.0
0.0
88.5
0.0
3.0
40.0
6.5
0.0
Table 4: Optimised interventions find hidden pathways; a downstream on-manifold penalty redirects unrestricted DAS. For each DAS variant: success rate (succ., %) and, among its successes, the share that is pathway-dominated (path., %). G: DAS trained on the model with planted pathways. H: the same, restricted to the complement of the top-32 principal components of the site and of the clean-model DAS direction (illusion probe). Main runs: all sites, seed 0. Mitigation runs (GPT-2, L=2 , T=4 , 900 interventions; gender: seeds 0–1, SVA: seed 0) retrain the same variants without (G ′ , H ′ ) and with (J, I) the downstream on-manifold penalty ( λ=5 ).
Model
Task
#succ.
full patch
mean-diff.
DAS
DAS+rand.
ρ
H succ.
GPT-2
Gender
2720
0.0
0.0
0.8
0.0
-0.15
1.3
GPT-2
IOI
2174
0.0
0.0
0.3
0.0
0.21
0.0
GPT-2
SVA
2809
0.0
0.0
0.0
0.0
-0.13
0.5
Positive control: wrong-mechanism successes with planted pathways
GPT-2
Gender + silent
1099
85.0
GPT-2
Gender + hijack
1393
96.1
Table 5: Unmodified models. Percentage of successful interventions whose decision relies on out-of-regime downstream activity (HPC-L >1 , i.e. more than one natural median margin); mean Spearman ρ between site divergence (LPCA) and HPC-L; success rate (%) of restricted DAS (H). Positive control (bottom): the same HPC-L >1 rate among wrong-mechanism successes in models with planted pathways, all intervention types pooled.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Model
Task
Pathway
#pos
site LPCA
site Maha.
Alg. 1
Down-OM
HPC-L
GPT-2
Gender
silent
1099
0.44 [0.42, 0.46]
0.64 [0.63, 0.66]
0.61 [0.59, 0.63]
0.65 [0.63, 0.66]
0.70 [0.69, 0.72]
GPT-2
Gender
hijack
1393
0.41 [0.39, 0.42]
0.60 [0.58, 0.62]
0.55 [0.54, 0.57]
0.60 [0.59, 0.62]
0.77 [0.75, 0.78]
GPT-2
Gender
distrib.
832
0.43 [0.41, 0.45]
0.63 [0.61, 0.65]
0.57 [0.55, 0.59]
0.59 [0.57, 0.61]
0.68 [0.66, 0.70]
GPT-2
Gender
in-range
71
0.28 [0.21, 0.35]
0.37 [0.29, 0.44]
0.23 [0.18, 0.27]
0.38 [0.31, 0.45]
0.48 [0.40, 0.56]
GPT-2
SVA
silent
24
0.41 [0.33, 0.49]
0.78 [0.73, 0.83]
0.86 [0.78, 0.93]
0.92 [0.88, 0.95]
0.99 [0.99, 1.00]
GPT-2
SVA
hijack
14
0.46 [0.37, 0.56]
0.81 [0.77, 0.84]
0.85 [0.79, 0.92]
0.94 [0.92, 0.96]
0.99 [0.99, 1.00]
Appendix
Table 6: Wrong-mechanism successes: pooled AUROC (scores rank-normalised within each configuration) with 95% bootstrap confidence intervals over interventions (1,000 resamples).
Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks. To study the role of model components in task behaviour, their causal influence can be investigated through interventions. Prior work on model steering has largely focused on interventions along global directions in activation space, modeling task representations as approximately linear and additive. By studying interventions at the neuron level, we find substantial, neuron-specific nonlinear effects on model outputs that are not captured by current steering approaches. We introduce Distributed Sparse Interventions (DSI), an intervention approach that considers nonlinearities and interactions between neurons across layers to identify sparse sets of neurons that elicit task-relevant computations. Across a range of tasks, we demonstrate that DSI can activate task behaviour in instruction-tuned language models by localising and intervening on as few as 0.01% of neurons, highlighting the effectiveness of sparse, distributed interventions in the neuron basis. Additionally, adopting a set-based perspective enables computations over the identified neuron sets, offering insights into the roles of individual neurons by analysing their effects across tasks. Through sparse interventions, DSI enables fine-grained control over model behaviour, localisation of task-relevant neuron sets, and furthers our understanding of task composition.
Maximilian S. Ernst, Lorenz Linhardt, Aaron Peikert +1
Max Planck School of Cognition · Center for Lifespan Psychology Max Planck Institute for Human Development · Machine Learning Group Technische Universität Berlin +5
Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with the low-level computation realized by a neural network through counterfactual intervention analysis. Existing methods such as distributed alignment search (DAS) learn expressive subspace interventions, but the relevant neural site is unknown a priori, so finding a handle requires a computationally burdensome search over candidate sites. We introduce PLOT (Progressive Localization via Optimal Transport), a transport-based framework that localizes causal variables from the output effect geometry of abstract and neural interventions. PLOT fits an optimal transport coupling between abstract variables and candidate neural sites, yielding a global soft correspondence that can be calibrated into intervention handles. In simple settings, a single coupling over individual neurons suffices. In larger models, PLOT is applied progressively, moving from coarse sites such as tokens, timesteps, or layers to finer supports such as coordinate groups or PCA spans, and optionally guiding DAS based on the localized signal. Across experiments of increasing complexity, transport-only PLOT handles are exceedingly fast and competitive on accuracy, while PLOT-guided DAS reaches DAS-level accuracy at a fraction of full DAS runtime, providing an efficient localization engine for causal abstraction research at scale.
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. In practice, however, feature entanglement leads to interference such that localized interventions can have unintended downstream effects. Motivated by the \textit{Independent Causal Mechanisms} principle, we propose to constrain internal features to be almost orthogonal. We argue that this promotes modular representations amenable to causal intervention. We formalize this problem by characterizing the gap between an idealized isolated intervention and its realized effect on model outputs in terms of feature interference. We upper-bound the propagation of feature interference in terms of the self-coherence of the feature dictionary, and relate this discrepancy to an explicit orthogonality regularization on the dictionary itself. Empirically, we show that this regularization enables more isolated interventions on mathematical reasoning concepts while preserving model performance. Our code is available under \texttt{https://github.com/mrtzmllr/sae-icm}.
Moritz Miller, Florent Draye, Bernhard Schölkopf
Max Planck Institute for Intelligent Systems, Tübingen, Germany · ETH Zurich, Switzerland · Max Planck ETH Center for Learning Systems +1