Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.
Figures & tables
Figure 1: Preservation of correct and incorrect model prediction across circuit sizes. Circuits are extracted on the full public-training split and evaluated on the public-test split. Shading indicate the point-wise 95% bootstrap confidence intervals. Specific models are: Llama-3.1 (8B), Gemma-2 (2b), Qwen-2.5 (0.5B).
Figure 2: Two omitted computations recover model errors: A) Restored contextual computation changes S-Inhibition and Name Mover queries, producing Courtney instead of Sara. B) Restored-name source computations changes keys, reducing the retrieval of Katie by a Positive Name Mover and increasing it by a Negative Name Mover, both favouring Vanessa over Katie.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Split
Rows evaluated
Candidate errors
Candidate correct
IOI
Discovery pool
13,000
135
12,865
IOI
Validation
10,000
95
9,905
IOI
Test census
100,000
884
99,116
IOI
Circuit-scored test sample
10,884
884
10,000
Docstring
Discovery pool
3,000
1,027
1,973
Docstring
Validation
2,000
676
1,324
Appendix
Table 2: IOI/Docstring data coverage. Discovery pools are used in full for current EAP-IG, Edge Pruning where supported, and OA fitting.
Task
Circuit
Epochs
Updates
Selected epoch
IOI
Manual / ACDC / EAP-IG
13
10,725
12
IOI
Edge Pruning seed 0
12
9,900
9
IOI
Edge Pruning seeds 1, 2
13
10,725
12
Docstring
Manual
26
4,888
23
Docstring
ACDC τ=.005
26
4,888
23
Docstring
ACDC τ=.02
29
5,452
26
Appendix
Table 3: OA training for current recovery/reference rows. Selection uses minimum validation KL
Approach
Observed result
Tradeoff or limitation
Full-training discovery
Correct–error agreement gaps persist on Gemma ARC-Challenge despite errors comprising 38.4% of discovery examples.
Greater discovery coverage has mixed effects across settings; exposure to errors does not ensure their reproduction.
Error reweighting
IOI ACDC: assigning errors 50% of discovery KL weight raises Aerr from 9.3% to 37.0%.
Aok falls from 98.7% to 91.0%.
Distinct-error enrichment
IOI ACDC: Aerr rises from 9.3% to 44.2% with 500 error and 500 correct discovery examples.
Aok falls from 98.7% to 91.4%; enrichment changes discovery composition as well as error prevalence.
Optimal ablation
Improves both agreement measures for some Docstring circuits; large IOI agreement gaps remain.
Benefits depend on the circuit and task; minimizing KL does not ensure high exact-error agreement.
Error-conditioned replacements
Mean and resample replacements increase IOI exact-error agreement.
Scalar subject-logit shifts match or exceed recovery at approximately matched correct-case cost. These controls are calibrated on evaluation correct cases.
Scalar and threshold calibration
Can substantially increase IOI exact-error agreement; matching overall error rates does not consistently recover particular errors.
Aggressive thresholds can narrow the agreement gap by reducing correct preservation.
Appendix
Table 4: Summary of attempts to improve exact reproduction of model errors, with their correct-case effects and limitations.
Figure 3: Correct minus error agreement gaps on MIB before and after margin matching. Circuits are extracted using EAP-IG on the full public-training split and evaluated on the public-test split. Correct and error examples are bucketed within each model–task setting within a 0.1 logit range without replacement; Shading denotes pointwise 95% bootstrap intervals.
The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by measuring necessity and sufficiency. But these criteria say little about whether a circuit consistently captures how a model performs a task, or if it is specific to that task. We study these two properties, consistency and specificity, across six tasks and five models, extracting circuits at the component level (attention heads and MLP blocks) and at the level of individual MLP neurons. We find that component-level circuits are highly consistent and causally important on most tasks, but they are not specific: ablating one task's circuit damages another task's performance about as much as that task's own circuit does. Neuron-level circuits, on the other hand, exhibit higher task-specificity but are far less consistent within tasks. This is explained by circuit overlap: component-level circuits share most of their components across all task pairs, related or not, while neuron-level circuits overlap only between closely related tasks. In a case study of the components shared by the task circuits of Llama-3.2-3B, we show that they consist mostly of MLP blocks, while the few attention heads within turn out to be generic attention-sink heads. Overall, our findings raise questions about the degree to which circuits can support targeted understanding of, and intervention on, model behavior.
Michael Li, Nishant Subramani
Carnegie Mellon University · Work done while at Carnegie Mellon University. · Northeastern University +1
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize. In this work, we study whether language model (LM) agents can assist with this explanation problem once a circuit has already been identified. We introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer circuits with 163 component-level annotations. We propose HyVE (Hypothesize, Validate, Explain), an agentic explainer that analyzes each component through an iterative loop of observation, hypothesis generation, and causal validation, eventually producing a component-level explanation and a circuit-level task description. Across four LM backbones, HyVE recovers useful component- and task-level explanations, but no backbone is uniformly best. Our analysis shows that strong backbones usually form observation-grounded hypotheses, while failures more often arise later in the validation loop, through incomplete validation plans, code execution errors, or unresolved hypotheses. A case study on an arithmetic circuit in Llama-3-8B shows that the same formulation can extend beyond semi-synthetic benchmarks to naturally trained models. Overall, LM agents are promising circuit explainers, but reliable validation remains the key obstacle.
Ayan Antik Khan, Harsh Kohli, Yuekun Yao +2
George Mason University · The Ohio State University