Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
Organizations: University of Toronto · McGill University · Tsinghua University
Abstract
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.
Figures & tables
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Split | Rows evaluated | Candidate errors | Candidate correct |
|---|---|---|---|---|
| IOI | Discovery pool | 13,000 | 135 | 12,865 |
| IOI | Validation | 10,000 | 95 | 9,905 |
| IOI | Test census | 100,000 | 884 | 99,116 |
| IOI | Circuit-scored test sample | 10,884 | 884 | 10,000 |
| Docstring | Discovery pool | 3,000 | 1,027 | 1,973 |
| Docstring | Validation | 2,000 | 676 | 1,324 |
| Task | Circuit | Epochs | Updates | Selected epoch |
|---|---|---|---|---|
| IOI | Manual / ACDC / EAP-IG | 13 | 10,725 | 12 |
| IOI | Edge Pruning seed 0 | 12 | 9,900 | 9 |
| IOI | Edge Pruning seeds 1, 2 | 13 | 10,725 | 12 |
| Docstring | Manual | 26 | 4,888 | 23 |
| Docstring | ACDC | 26 | 4,888 | 23 |
| Docstring | ACDC | 29 | 5,452 | 26 |
| Approach | Observed result | Tradeoff or limitation |
|---|---|---|
| Full-training discovery | Correct–error agreement gaps persist on Gemma ARC-Challenge despite errors comprising 38.4% of discovery examples. | Greater discovery coverage has mixed effects across settings; exposure to errors does not ensure their reproduction. |
| Error reweighting | IOI ACDC: assigning errors 50% of discovery KL weight raises from 9.3% to 37.0%. | falls from 98.7% to 91.0%. |
| Distinct-error enrichment | IOI ACDC: rises from 9.3% to 44.2% with 500 error and 500 correct discovery examples. | falls from 98.7% to 91.4%; enrichment changes discovery composition as well as error prevalence. |
| Optimal ablation | Improves both agreement measures for some Docstring circuits; large IOI agreement gaps remain. | Benefits depend on the circuit and task; minimizing KL does not ensure high exact-error agreement. |
| Error-conditioned replacements | Mean and resample replacements increase IOI exact-error agreement. | Scalar subject-logit shifts match or exceed recovery at approximately matched correct-case cost. These controls are calibrated on evaluation correct cases. |
| Scalar and threshold calibration | Can substantially increase IOI exact-error agreement; matching overall error rates does not consistently recover particular errors. | Aggressive thresholds can narrow the agreement gap by reducing correct preservation. |
| Model | Task | CPR | CMD |
|---|---|---|---|
| Llama-3.1 8B | ARC-Challenge | 0.8388 | 0.1602 |
| Gemma-2 2B | ARC-Challenge | 1.0017 | 0.0547 |
| Qwen-2.5 0.5B | ARC-Challenge | 0.9364 | 0.0646 |
| Llama-3.1 8B | ARC-Easy | 0.8458 | 0.1639 |
| Gemma-2 2B | ARC-Easy | 0.9943 | 0.0392 |
| Llama-3.1 8B | Subtraction | 0.9954 | 0.0036 |
| Reference group | Heads | Retained position |
|---|---|---|
| Name Movers | END | |
| Backup Name Movers | END | |
| Negative Name Movers | END | |
| S-Inhibition | END | |
| Induction | S2 | |
| Duplicate Token | S2 |