Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
Figures & tables
Figure 1: The objective-level recovery gap. Replacing excluded signals with donor activations, means, or zeros changes a circuit’s computational context without changing its structure. This context distortion can make faithfulness prefer B even when an equally sized A better reproduces intact-model behavior under a fixed behavioral test.
Task
Behavioral essence
Source
IOI
Indirect-object identification: recover the name that fills the repeated syntactic role.
( Wang et al., 2023 )
Greater-Than
Numerical comparison: determine whether one two-digit year ending is strictly greater than another.
( Hanna et al., 2023 )
Docstring
Documentation retrieval: predict the token sequence associated with a function’s docstring behavior.
( Heimersheim and Janiak, 2023 )
Acronym
Acronym completion: map a multiword description to its abbreviated form.
( García-Carrasco et al., 2024 )
InterpBench suite
Ten semi-synthetic tasks with native-closure-verified references: 113, 97, 2, 82, 111, 45, 58, 93, 103, and 25.
( Gupta et al., 2024 )
Table 1: The benchmark suites used for controlled and discovery comparisons. The first four rows are human circuits; the final row is the corrected ten-task InterpBench panel.
Task
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
Human suite
IOI
6.7 (0.33)
7.3 (3.56)
33.3 (8.82)
23.4 (8.05)
3.3 (0.50)
21.1 (15.90)
Greater-Than
0.0 (–)
3.4 (1.85)
0.0 (–)
5.9 (1.28)
56.7 (28.87)
54.0 (27.20)
Docstring
8.0 (0.92)
3.3 (0.99)
14.0 (4.86)
12.9 (5.00)
32.0 (4.40)
16.0 (6.03)
Acronym
5.0 (0.28)
5.8 (3.94)
15.0 (11.50)
20.6 (14.44)
10.0 (18.36)
29.2 (20.46)
Table 2: KL misranking. Each cell reports the percentage of total pairs that misrank, followed in parentheses by mean ΔQ (percentage points) among those misrankings for that intervention and pair type. Q ties and score ties count as non-misranking outcomes.
Figure 2: Edit magnitude does not determine misranking monotonically. Columns show human-reference tasks; rows show R–C and within-band C–C comparisons. The y-axis is the KL misranking rate over all pairs. Colors and markers denote interventions. The first band changes one component. Repeated points for merged bands reuse the same candidate set. Overall C–C results in Table 2 also include between-band pairs. Exact denominators are in Appendix B.4 .
Metric
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
KL
3.0 (0.36)
3.7 (3.03)
8.8 (4.28)
9.3 (8.31)
29.5 (28.83)
33.3 (28.03)
LD
6.1 (1.07)
6.2 (2.70)
6.4 (3.22)
7.7 (5.68)
18.1 (13.05)
19.0 (16.37)
PD
6.5 (12.13)
5.0 (9.16)
6.1 (1.33)
8.8 (8.72)
25.1 (20.73)
25.1 (20.39)
Table 3: KL/LD/PD on common pair identities across the 14 controlled tasks. Each cell reports percentage misranking followed by mean ΔQ in percentage points for that metric, intervention, and pair type; the denominator is the total common-pair pool.
Method
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
EAP
13.8 (1.92)
20.0 (2.18)
48.3 (6.01)
31.0 (3.67)
20.7 (5.17)
40.0 (3.76)
EAP-IG
35.5 (1.33)
30.4 (1.36)
38.7 (1.17)
53.0 (1.49)
29.0 (0.67)
45.2 (1.25)
ACDC
57.1 (1.67)
41.2 (1.11)
50.0 (1.76)
39.2 (1.47)
28.6 (1.83)
43.1 (1.86)
Edge-SP
0.0 (–)
9.4 (1.73)
20.0 (7.92)
14.4 (1.91)
17.5 (34.14)
31.1 (8.42)
Table 4: Human suite discovery: KL misranking among primary size-controlled pairs. Each cell reports percentage misranking followed by mean ΔQ in percentage points for that intervention and pair type; Q ties count as non-misranking outcomes.
Method
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
EAP
5.0 (0.68)
8.7 (0.14)
16.0 (0.66)
13.3 (0.79)
49.0 (1.56)
25.6 (5.34)
EAP-IG
0.0 (–)
2.9 (0.07)
6.1 (1.10)
2.9 (0.10)
46.5 (1.44)
5.0 (0.06)
ACDC
10.5 (0.48)
7.9 (0.43)
18.4 (0.63)
24.1 (1.88)
30.3 (2.48)
46.0 (2.10)
Edge-SP
0.0 (–)
18.7 (0.15)
6.0 (0.51)
27.1 (0.25)
40.0 (1.36)
37.6 (0.27)
Table 5: InterpBench suite discovery: KL misranking among primary size-controlled pairs. Each cell reports percentage misranking followed by mean ΔQ in percentage points for that intervention and pair type; Q ties count as non-misranking outcomes.
Confirmed repair rates (%)
Intervention
Pair
Mean ΔQ (pp) before (after)
80%
40%
20%
10%
Any level
Human suite (50 cases)
Resampling
R–C
1.50(−1.50)
100.0
100.0
50.0
50.0
100.0
C–C
4.67(−4.67)
87.0
95.7
87.0
91.3
100.0
Mean
R–C
14.21(−14.21)
100.0
100.0
80.0
80.0
100.0
C–C
25.96(−25.96)
66.7
100.0
100.0
100.0
100.0
Table 6: Context restoration repairs persistent misrankings. Entries are independent-test-confirmed repair rates in percent. Columns 80%–10% denote nominal restoration levels; successes can overlap across levels. Any level counts each case once. Mean ΔQ is shown as before (after) in pp, with signed deficit Q(nonpreferred)−Q(preferred) . Before means include all selected cases; after means include only confirmed repairs, once per case. The underlying Q values remain fixed.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Reference
1 edit
5%
10%
20%
50%
75%
IOI
99.00
99.00 (10)
97.80 (10)
83.33 (10)
59.40 (10)
58.20 (10)
49.30 (10)
Greater-Than
99.67
99.33 (10)
97.67 (10)
85.03 (10)
84.47 (10)
55.00 (10)
53.70 (10)
Docstring
34.00
32.57 (10)
32.57 (10)
28.20 (10)
25.03 (10)
12.63 (10)
9.67 (10)
Acronym
87.67
80.67 (10)
80.67 (10)
80.67 (10)
70.66 (10)
43.92 (10)
23.51 (10)
Appendix
Table 9: Section 3 Human suite: reference Q and mean candidate Q (percent), with unique candidate count in parentheses per band. Merged bands share the same candidates; those are not duplicated in pooled results. Greater-Than uses semantic accuracy; other tasks use intact-model answer agreement. All Q evaluations use ordinary resampling.
Task
Reference
1 edit
5%
10%
20%
50%
75%
113
99.93
99.92 (10)
96.59 (10)
92.73 (10)
60.46 (10)
3.59 (10)
3.11 (10)
97
100.00
100.00 (10)
75.40 (10)
44.49 (10)
25.01 (10)
19.49 (10)
7.37 (10)
2
97.78
92.18 (10)
84.39 (10)
60.45 (10)
53.60 (10)
3.50 (10)
3.48 (10)
82
99.44
97.07 (10)
67.29 (10)
65.94 (10)
40.41 (10)
15.60 (10)
10.22 (10)
111
97.33
96.79 (10)
94.75 (10)
92.88 (10)
92.93 (10)
90.35 (10)
90.19 (10)
45
98.22
98.17 (10)
69.28 (10)
55.97 (10)
38.63 (10)
10.58 (10)
10.62 (10)
Appendix
Table 10: Section 3 InterpBench suite: reference Q and mean candidate Q (percent), with unique candidate count in parentheses per band. Merged bands share the same candidates; those are not duplicated in pooled results. Greater-Than uses semantic accuracy; other tasks use intact-model answer agreement. All Q evaluations use ordinary resampling.
Task
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
Human suite
IOI
4/60
130/1770
20/60
414/1770
2/60
373/1770
Greater-Than
0/60
61/1770
0/60
104/1770
34/60
956/1770
Docstring
4/50
41/1225
7/50
158/1225
16/50
196/1225
Acronym
2/40
45/780
6/40
161/780
4/40
228/780
Appendix
Table 11: Complete Section 3 KL misranking counts. Every cell is misranked / total pairs, with Q ties included in the denominator.
Task
Band
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
IOI
1 edit
2/10
15/45
1/10
10/45
0/10
1/45
IOI
5%
2/10
8/45
7/10
34/45
2/10
15/45
IOI
10%
0/10
4/45
8/10
10/45
0/10
21/45
IOI
20%
0/10
12/45
4/10
6/45
0/10
13/45
IOI
50%
0/10
4/45
0/10
4/45
0/10
29/45
Appendix
Table 12: Exact misranked / total pair counts underlying Figure 2 . C–C comparisons in this table are within the same band; between-band pairs appear in the overall task table. Tied pairs remain in the denominator.
Task
Metric
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
IOI
LD
1/60
49/1770
2/60
121/1770
46/60
1406/1770
IOI
PD
2/60
172/1770
2/60
105/1770
45/60
1402/1770
Greater-Than
LD
1/60
159/1770
1/60
180/1770
6/60
398/1770
Greater-Than
PD
0/60
43/1770
0/60
77/1770
12/60
479/1770
Docstring
LD
4/50
44/1225
9/50
175/1225
8/50
167/1225
Appendix
Table 13: Section 3 task-level LD and PD misranking counts. Each entry is misranked pairs / total pairs within its own fixed task universe, including Q and score ties in the denominator. A dash means no primary pair pool. The pooled common-metric comparison uses the intersection of pair identities rather than pooling these denominators.
Task
Reference
EAP
EAP-IG
ACDC
Edge-SP
IOI
99.00
90.70 (10)
98.10 (10)
98.50 (10)
58.80 (10)
Greater-Than
99.67
96.38 (8)
99.48 (9)
100.00 (4)
59.87 (10)
Docstring
34.00
30.37 (9)
34.47 (10)
34.90 (10)
35.57 (10)
Acronym
87.67
84.53 (4)
83.00 (4)
85.17 (4)
23.80 (10)
Appendix
Table 15: Section 4 Human suite: reference Q and mean saved-candidate Q (percent), with candidate count in parentheses. This performance inventory includes non-primary sizes, which are excluded from the primary misranking tables. Methods are not ranked by this table. Greater-Than uses semantic accuracy; other tasks use intact-model answer agreement. All Q evaluations use ordinary resampling.
Task
Reference
EAP
EAP-IG
ACDC
Edge-SP
113
99.93
66.51 (10)
100.00 (10)
97.30 (10)
96.83 (10)
97
100.00
25.90 (10)
100.00 (10)
97.62 (4)
100.00 (10)
2
97.78
100.00 (10)
100.00 (10)
98.14 (6)
99.34 (10)
82
99.44
80.53 (10)
100.00 (10)
94.71 (10)
96.03 (10)
111
97.33
96.99 (10)
99.85 (10)
98.01 (10)
99.01 (10)
45
98.22
98.32 (10)
99.75 (10)
97.15 (10)
93.96 (10)
Appendix
Table 16: Section 4 InterpBench suite: reference Q and mean saved-candidate Q (percent), with candidate count in parentheses. This performance inventory includes non-primary sizes, which are excluded from the primary misranking tables. Methods are not ranked by this table. Greater-Than uses semantic accuracy; other tasks use intact-model answer agreement. All Q evaluations use ordinary resampling.
Task
Method
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
IOI
EAP
0/10
11/45
9/10
22/45
0/10
14/45
EAP-IG
7/10
15/45
7/10
31/45
2/10
23/45
ACDC
—
—
—
—
—
—
Edge-SP
0/10
8/45
1/10
8/45
0/10
16/45
Greater-Than
EAP
1/8
4/28
1/8
3/28
3/8
13/28
Appendix
Table 17: Section 4 task-level KL misranking counts (misranked/total pairs). Ties count as non-misranking outcomes; dashes denote unavailable pair pools.
Task
Method
Metric
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
IOI
EAP
LD
0/10
5/45
0/10
21/45
10/10
25/45
IOI
EAP
PD
0/10
13/45
0/10
23/45
10/10
23/45
IOI
EAP-IG
LD
2/10
9/45
2/10
10/45
7/10
19/45
IOI
EAP-IG
PD
2/10
10/45
2/10
11/45
7/10
20/45
IOI
ACDC
LD
–
–
–
–
–
–
Appendix
Table 18: Section 4 task-level LD and PD misranking counts. Each entry is misranked pairs / total pairs within its own fixed task and method universe, including Q and score ties in the denominator. A dash means no primary pair pool. The pooled common-metric comparison uses the intersection of pair identities rather than pooling these denominators.
Metric
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
KL
20.2 (1.55)
20.9 (1.55)
36.0 (4.24)
30.9 (2.05)
22.8 (10.90)
38.1 (4.28)
LD
12.3 (1.79)
17.9 (1.43)
23.7 (3.22)
28.0 (2.69)
43.9 (11.86)
45.7 (6.40)
PD
19.3 (3.46)
19.5 (1.72)
20.2 (2.93)
35.2 (4.72)
53.5 (10.16)
52.9 (7.09)
Appendix
Table 20: KL/LD/PD on common pair identities in the Human suite. Each cell reports percentage misranking followed by mean ΔQ in percentage points for that metric, intervention, and pair type; the denominator is the total common-pair pool.
Metric
Resampling
Mean
Zero
R–C
C–C
R–C
C–C
R–C
C–C
KL
3.5 (0.56)
9.8 (0.18)
11.2 (0.69)
16.2 (0.79)
42.1 (1.61)
26.8 (2.14)
LD
26.4 (1.95)
20.9 (0.59)
24.3 (1.90)
22.5 (2.09)
45.6 (1.36)
23.8 (2.53)
PD
2.7 (0.38)
8.2 (0.17)
18.7 (1.18)
15.8 (0.80)
48.5 (1.45)
27.5 (2.22)
Appendix
Table 21: KL/LD/PD on common pair identities in the InterpBench suite. Each cell reports percentage misranking followed by mean ΔQ in percentage points for that metric, intervention, and pair type; the denominator is the total common-pair pool.
Confirmed repairs ( n/N )
Intervention
Pair
Mean ΔQ (pp) before (after)
80%
40%
20%
10%
Any level
Human suite (50 cases)
Resampling
R–C
1.50(−1.50)
2/2
2/2
1/2
1/2
2/2
C–C
4.67(−4.67)
20/23
22/23
20/23
21/23
23/23
Mean
R–C
14.21(−14.21)
10/10
10/10
8/10
8/10
10/10
C–C
25.96(−25.96)
2/3
3/3
3/3
3/3
3/3
Appendix
Table 24: Independent-test-confirmed repairs as n/N , with the same rows and denominators as Table 6 . Columns 80%–10% denote nominal restoration levels; successes can overlap across levels. Any level counts each case once. Mean ΔQ is shown as before (after) in pp, with signed deficit Q(nonpreferred)−Q(preferred) . Before means include all selected cases; after means include only confirmed repairs, once per case. The underlying Q values remain fixed.
Task
Intervention
Pair type
Baseline ΔQ (pp)
After-repair ΔQ (pp)
80
40
20
10
Human suite
Acronym
Mean
R–C
15.44
-15.44
Y
Y
Y
Y
Acronym
Resampling
C–C
11.33
-11.33
Y
Y
Y
Y
Acronym
Resampling
C–C
0.56
-0.56
Y
Y
Y
Y
Acronym
Resampling
C–C
12.33
-12.33
Y
Y
Y
Y
Acronym
Resampling
C–C
7.44
-7.44
Y
Y
N
Y
Appendix
Table 25: Complete Section 5 per-case catalog. Gaps use the same score-preference orientation as Table 6 ; Q values remain fixed. Y/N indicate independent-test confirmation. A dash in the after-repair column denotes no confirmed repair, rather than a negative gap for a failed case.
Stratum
Any level
80%
40%
20%
10%
Human suite
50/50
45/50
48/50
44/50
43/50
InterpBench suite
46/50
39/50
40/50
43/50
39/50
Resampling
50/50
44/50
49/50
45/50
43/50
Mean
23/25
21/25
21/25
20/25
19/25
Zero
23/25
19/25
18/25
22/25
20/25
R–C
46/48
44/48
42/48
40/48
37/48
Appendix
Table 26: Section 5 100-case cohort breakdown. Entries are confirmed successes / selected cases; levels are non-exclusive.
Task
Intervention
Pair
KL gap
LD gap
PD gap
Q gap
IOI
mean
k193_s810 / k193_s816
0.1192
0.1585
0.0006241
2
Greater-Than
mean
k12_s811 / k12_s813
0.07244
0.03427
0.01234
2
docstring
mean
k12_s811 / k12_s814
0.06245
0.2988
0.027
2
acronym
mean
k2_s818 / k6_s812
0.02785
0.4786
0.1154
8.778
Appendix
Table 29: Shared-failure examples: lexicographically first all-three failure per Human suite task (no gap filter). Score gaps favor the worse-Q candidate. Q gaps are in percentage points.
Panel / pair
Shared
Median [Q1, Q3] (pp)
≥5 pp
Joint support
S3 Human suite R–C
15
2.33 [0.50, 4.67]
4
2/13
S3 Human suite C–C
597
3.67 [1.00, 18.67]
271
152/472
S3 InterpBench suite R–C
60
2.19 [0.16, 7.15]
25
38/60
S3 InterpBench suite C–C
1784
4.31 [0.74, 6.97]
799
1158/1784
S4 Human suite R–C
22
1.00 [0.33, 4.83]
6
5/16
S4 Human suite C–C
128
1.00 [0.67, 2.33]
12
13/106
Appendix
Table 30: Agreement deficits among all-three-metric failures. Joint support: supported / available saved paired-support records, requiring support for all three metrics. A dash means no complete joint-uncertainty record; it does not mean zero supported failures.
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.
Li Zhang, Chuqin Geng, Mark Zhang +4
University of Toronto · McGill University · Tsinghua University
Mechanistic interpretability seeks to explain a model's behaviour by finding its circuit: the sparse subgraph of the model's computation that is causally responsible for it. Automated methods have made this search systematic, but each one starts afresh for every behaviour, and the effort spent finding one circuit does nothing for the next. Circuit discovery has thus been automated, but not amortised. We ask whether circuit discovery can itself be learned. We frame it as a sequential decision problem over the computation graph of GPT-2 small, in which a policy removes edges until it reaches a compact subgraph that preserves the behaviour, guided by a faithfulness reward defined through causal intervention. A single policy trained across twelve behaviours recovers a faithful circuit for each, and once frozen it transfers to behaviours it never saw during training, recovering their known circuits without further search. A short warm-start improves these transferred circuits, returning far smaller ones than training from scratch. While the learned policy does not match a per-behaviour search on circuit size or cost, it shows that circuit discovery is a learnable, transferable procedure rather than a search repeated for every behaviour.
Barsat Khadka
School of Computing Sciences and Engineering The University of Southern Mississippi
The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by measuring necessity and sufficiency. But these criteria say little about whether a circuit consistently captures how a model performs a task, or if it is specific to that task. We study these two properties, consistency and specificity, across six tasks and five models, extracting circuits at the component level (attention heads and MLP blocks) and at the level of individual MLP neurons. We find that component-level circuits are highly consistent and causally important on most tasks, but they are not specific: ablating one task's circuit damages another task's performance about as much as that task's own circuit does. Neuron-level circuits, on the other hand, exhibit higher task-specificity but are far less consistent within tasks. This is explained by circuit overlap: component-level circuits share most of their components across all task pairs, related or not, while neuron-level circuits overlap only between closely related tasks. In a case study of the components shared by the task circuits of Llama-3.2-3B, we show that they consist mostly of MLP blocks, while the few attention heads within turn out to be generic attention-sink heads. Overall, our findings raise questions about the degree to which circuits can support targeted understanding of, and intervention on, model behavior.
Michael Li, Nishant Subramani
Carnegie Mellon University · Work done while at Carnegie Mellon University. · Northeastern University +1