Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability
Organizations: University of Toronto · McGill University
Abstract
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
Figures & tables
| Task | Behavioral essence | Source |
| IOI | Indirect-object identification: recover the name that fills the repeated syntactic role. | ( Wang et al., 2023 ) |
| Greater-Than | Numerical comparison: determine whether one two-digit year ending is strictly greater than another. | ( Hanna et al., 2023 ) |
| Docstring | Documentation retrieval: predict the token sequence associated with a function’s docstring behavior. | ( Heimersheim and Janiak, 2023 ) |
| Acronym | Acronym completion: map a multiword description to its abbreviated form. | ( García-Carrasco et al., 2024 ) |
| InterpBench suite | Ten semi-synthetic tasks with native-closure-verified references: 113, 97, 2, 82, 111, 45, 58, 93, 103, and 25. | ( Gupta et al., 2024 ) |
| Task | Resampling | Mean | Zero | |||
| R–C | C–C | R–C | C–C | R–C | C–C | |
| Human suite | ||||||
| IOI | 6.7 (0.33) | 7.3 (3.56) | 33.3 (8.82) | 23.4 (8.05) | 3.3 (0.50) | 21.1 (15.90) |
| Greater-Than | 0.0 (–) | 3.4 (1.85) | 0.0 (–) | 5.9 (1.28) | 56.7 (28.87) | 54.0 (27.20) |
| Docstring | 8.0 (0.92) | 3.3 (0.99) | 14.0 (4.86) | 12.9 (5.00) | 32.0 (4.40) | 16.0 (6.03) |
| Acronym | 5.0 (0.28) | 5.8 (3.94) | 15.0 (11.50) | 20.6 (14.44) | 10.0 (18.36) | 29.2 (20.46) |
| Metric | Resampling | Mean | Zero | |||
| R–C | C–C | R–C | C–C | R–C | C–C | |
| KL | 3.0 (0.36) | 3.7 (3.03) | 8.8 (4.28) | 9.3 (8.31) | 29.5 (28.83) | 33.3 (28.03) |
| LD | 6.1 (1.07) | 6.2 (2.70) | 6.4 (3.22) | 7.7 (5.68) | 18.1 (13.05) | 19.0 (16.37) |
| PD | 6.5 (12.13) | 5.0 (9.16) | 6.1 (1.33) | 8.8 (8.72) | 25.1 (20.73) | 25.1 (20.39) |
| Method | Resampling | Mean | Zero | |||
| R–C | C–C | R–C | C–C | R–C | C–C | |
| EAP | 13.8 (1.92) | 20.0 (2.18) | 48.3 (6.01) | 31.0 (3.67) | 20.7 (5.17) | 40.0 (3.76) |
| EAP-IG | 35.5 (1.33) | 30.4 (1.36) | 38.7 (1.17) | 53.0 (1.49) | 29.0 (0.67) | 45.2 (1.25) |
| ACDC | 57.1 (1.67) | 41.2 (1.11) | 50.0 (1.76) | 39.2 (1.47) | 28.6 (1.83) | 43.1 (1.86) |
| Edge-SP | 0.0 (–) | 9.4 (1.73) | 20.0 (7.92) | 14.4 (1.91) | 17.5 (34.14) | 31.1 (8.42) |
| Method | Resampling | Mean | Zero | |||
| R–C | C–C | R–C | C–C | R–C | C–C | |
| EAP | 5.0 (0.68) | 8.7 (0.14) | 16.0 (0.66) | 13.3 (0.79) | 49.0 (1.56) | 25.6 (5.34) |
| EAP-IG | 0.0 (–) | 2.9 (0.07) | 6.1 (1.10) | 2.9 (0.10) | 46.5 (1.44) | 5.0 (0.06) |
| ACDC | 10.5 (0.48) | 7.9 (0.43) | 18.4 (0.63) | 24.1 (1.88) | 30.3 (2.48) | 46.0 (2.10) |
| Edge-SP | 0.0 (–) | 18.7 (0.15) | 6.0 (0.51) | 27.1 (0.25) | 40.0 (1.36) | 37.6 (0.27) |
| Confirmed repair rates (%) | |||||||
| Intervention | Pair | Mean (pp) before (after) | 80% | 40% | 20% | 10% | Any level |
| Human suite (50 cases) | |||||||
| Resampling | R–C | 100.0 | 100.0 | 50.0 | 50.0 | 100.0 | |
| C–C | 87.0 | 95.7 | 87.0 | 91.3 | 100.0 | ||
| Mean | R–C | 100.0 | 100.0 | 80.0 | 80.0 | 100.0 | |
| C–C | 66.7 | 100.0 | 100.0 | 100.0 | 100.0 | ||
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Reference | 1 edit | 5% | 10% | 20% | 50% | 75% |
| IOI | 99.00 | 99.00 (10) | 97.80 (10) | 83.33 (10) | 59.40 (10) | 58.20 (10) | 49.30 (10) |
| Greater-Than | 99.67 | 99.33 (10) | 97.67 (10) | 85.03 (10) | 84.47 (10) | 55.00 (10) | 53.70 (10) |
| Docstring | 34.00 | 32.57 (10) | 32.57 (10) | 28.20 (10) | 25.03 (10) | 12.63 (10) | 9.67 (10) |
| Acronym | 87.67 | 80.67 (10) | 80.67 (10) | 80.67 (10) | 70.66 (10) | 43.92 (10) | 23.51 (10) |
| Task | Reference | 1 edit | 5% | 10% | 20% | 50% | 75% |
| 113 | 99.93 | 99.92 (10) | 96.59 (10) | 92.73 (10) | 60.46 (10) | 3.59 (10) | 3.11 (10) |
| 97 | 100.00 | 100.00 (10) | 75.40 (10) | 44.49 (10) | 25.01 (10) | 19.49 (10) | 7.37 (10) |
| 2 | 97.78 | 92.18 (10) | 84.39 (10) | 60.45 (10) | 53.60 (10) | 3.50 (10) | 3.48 (10) |
| 82 | 99.44 | 97.07 (10) | 67.29 (10) | 65.94 (10) | 40.41 (10) | 15.60 (10) | 10.22 (10) |
| 111 | 97.33 | 96.79 (10) | 94.75 (10) | 92.88 (10) | 92.93 (10) | 90.35 (10) | 90.19 (10) |
| 45 | 98.22 | 98.17 (10) | 69.28 (10) | 55.97 (10) | 38.63 (10) | 10.58 (10) | 10.62 (10) |
| Task | Resampling | Mean | Zero | |||
| R–C | C–C | R–C | C–C | R–C | C–C | |
| Human suite | ||||||
| IOI | 4/60 | 130/1770 | 20/60 | 414/1770 | 2/60 | 373/1770 |
| Greater-Than | 0/60 | 61/1770 | 0/60 | 104/1770 | 34/60 | 956/1770 |
| Docstring | 4/50 | 41/1225 | 7/50 | 158/1225 | 16/50 | 196/1225 |
| Acronym | 2/40 | 45/780 | 6/40 | 161/780 | 4/40 | 228/780 |
| Task | Band | Resampling | Mean | Zero | |||
|---|---|---|---|---|---|---|---|
| R–C | C–C | R–C | C–C | R–C | C–C | ||
| IOI | 1 edit | 2/10 | 15/45 | 1/10 | 10/45 | 0/10 | 1/45 |
| IOI | 5% | 2/10 | 8/45 | 7/10 | 34/45 | 2/10 | 15/45 |
| IOI | 10% | 0/10 | 4/45 | 8/10 | 10/45 | 0/10 | 21/45 |
| IOI | 20% | 0/10 | 12/45 | 4/10 | 6/45 | 0/10 | 13/45 |
| IOI | 50% | 0/10 | 4/45 | 0/10 | 4/45 | 0/10 | 29/45 |
| Task | Metric | Resampling | Mean | Zero | |||
|---|---|---|---|---|---|---|---|
| R–C | C–C | R–C | C–C | R–C | C–C | ||
| IOI | LD | 1/60 | 49/1770 | 2/60 | 121/1770 | 46/60 | 1406/1770 |
| IOI | PD | 2/60 | 172/1770 | 2/60 | 105/1770 | 45/60 | 1402/1770 |
| Greater-Than | LD | 1/60 | 159/1770 | 1/60 | 180/1770 | 6/60 | 398/1770 |
| Greater-Than | PD | 0/60 | 43/1770 | 0/60 | 77/1770 | 12/60 | 479/1770 |
| Docstring | LD | 4/50 | 44/1225 | 9/50 | 175/1225 | 8/50 | 167/1225 |
| Task | Reference | EAP | EAP-IG | ACDC | Edge-SP |
| IOI | 99.00 | 90.70 (10) | 98.10 (10) | 98.50 (10) | 58.80 (10) |
| Greater-Than | 99.67 | 96.38 (8) | 99.48 (9) | 100.00 (4) | 59.87 (10) |
| Docstring | 34.00 | 30.37 (9) | 34.47 (10) | 34.90 (10) | 35.57 (10) |
| Acronym | 87.67 | 84.53 (4) | 83.00 (4) | 85.17 (4) | 23.80 (10) |
| Task | Reference | EAP | EAP-IG | ACDC | Edge-SP |
| 113 | 99.93 | 66.51 (10) | 100.00 (10) | 97.30 (10) | 96.83 (10) |
| 97 | 100.00 | 25.90 (10) | 100.00 (10) | 97.62 (4) | 100.00 (10) |
| 2 | 97.78 | 100.00 (10) | 100.00 (10) | 98.14 (6) | 99.34 (10) |
| 82 | 99.44 | 80.53 (10) | 100.00 (10) | 94.71 (10) | 96.03 (10) |
| 111 | 97.33 | 96.99 (10) | 99.85 (10) | 98.01 (10) | 99.01 (10) |
| 45 | 98.22 | 98.32 (10) | 99.75 (10) | 97.15 (10) | 93.96 (10) |
| Task | Method | Resampling | Mean | Zero | |||
| R–C | C–C | R–C | C–C | R–C | C–C | ||
| IOI | EAP | 0/10 | 11/45 | 9/10 | 22/45 | 0/10 | 14/45 |
| EAP-IG | 7/10 | 15/45 | 7/10 | 31/45 | 2/10 | 23/45 | |
| ACDC | — | — | — | — | — | — | |
| Edge-SP | 0/10 | 8/45 | 1/10 | 8/45 | 0/10 | 16/45 | |
| Greater-Than | EAP | 1/8 | 4/28 | 1/8 | 3/28 | 3/8 | 13/28 |
| Task | Method | Metric | Resampling | Mean | Zero | |||
|---|---|---|---|---|---|---|---|---|
| R–C | C–C | R–C | C–C | R–C | C–C | |||
| IOI | EAP | LD | 0/10 | 5/45 | 0/10 | 21/45 | 10/10 | 25/45 |
| IOI | EAP | PD | 0/10 | 13/45 | 0/10 | 23/45 | 10/10 | 23/45 |
| IOI | EAP-IG | LD | 2/10 | 9/45 | 2/10 | 10/45 | 7/10 | 19/45 |
| IOI | EAP-IG | PD | 2/10 | 10/45 | 2/10 | 11/45 | 7/10 | 20/45 |
| IOI | ACDC | LD | – | – | – | – | – | – |
| Metric | Resampling | Mean | Zero | |||
| R–C | C–C | R–C | C–C | R–C | C–C | |
| KL | 20.2 (1.55) | 20.9 (1.55) | 36.0 (4.24) | 30.9 (2.05) | 22.8 (10.90) | 38.1 (4.28) |
| LD | 12.3 (1.79) | 17.9 (1.43) | 23.7 (3.22) | 28.0 (2.69) | 43.9 (11.86) | 45.7 (6.40) |
| PD | 19.3 (3.46) | 19.5 (1.72) | 20.2 (2.93) | 35.2 (4.72) | 53.5 (10.16) | 52.9 (7.09) |
| Metric | Resampling | Mean | Zero | |||
| R–C | C–C | R–C | C–C | R–C | C–C | |
| KL | 3.5 (0.56) | 9.8 (0.18) | 11.2 (0.69) | 16.2 (0.79) | 42.1 (1.61) | 26.8 (2.14) |
| LD | 26.4 (1.95) | 20.9 (0.59) | 24.3 (1.90) | 22.5 (2.09) | 45.6 (1.36) | 23.8 (2.53) |
| PD | 2.7 (0.38) | 8.2 (0.17) | 18.7 (1.18) | 15.8 (0.80) | 48.5 (1.45) | 27.5 (2.22) |
| Confirmed repairs ( ) | |||||||
| Intervention | Pair | Mean (pp) before (after) | 80% | 40% | 20% | 10% | Any level |
| Human suite (50 cases) | |||||||
| Resampling | R–C | 2/2 | 2/2 | 1/2 | 1/2 | 2/2 | |
| C–C | 20/23 | 22/23 | 20/23 | 21/23 | 23/23 | ||
| Mean | R–C | 10/10 | 10/10 | 8/10 | 8/10 | 10/10 | |
| C–C | 2/3 | 3/3 | 3/3 | 3/3 | 3/3 | ||
| Task | Intervention | Pair type | Baseline (pp) | After-repair (pp) | 80 | 40 | 20 | 10 |
|---|---|---|---|---|---|---|---|---|
| Human suite | ||||||||
| Acronym | Mean | R–C | 15.44 | -15.44 | Y | Y | Y | Y |
| Acronym | Resampling | C–C | 11.33 | -11.33 | Y | Y | Y | Y |
| Acronym | Resampling | C–C | 0.56 | -0.56 | Y | Y | Y | Y |
| Acronym | Resampling | C–C | 12.33 | -12.33 | Y | Y | Y | Y |
| Acronym | Resampling | C–C | 7.44 | -7.44 | Y | Y | N | Y |
| Stratum | Any level | 80% | 40% | 20% | 10% |
| Human suite | 50/50 | 45/50 | 48/50 | 44/50 | 43/50 |
| InterpBench suite | 46/50 | 39/50 | 40/50 | 43/50 | 39/50 |
| Resampling | 50/50 | 44/50 | 49/50 | 45/50 | 43/50 |
| Mean | 23/25 | 21/25 | 21/25 | 20/25 | 19/25 |
| Zero | 23/25 | 19/25 | 18/25 | 22/25 | 20/25 |
| R–C | 46/48 | 44/48 | 42/48 | 40/48 | 37/48 |
| Task | Intervention | Pair | KL gap | LD gap | PD gap | Q gap |
| IOI | mean | k193_s810 / k193_s816 | 0.1192 | 0.1585 | 0.0006241 | 2 |
| Greater-Than | mean | k12_s811 / k12_s813 | 0.07244 | 0.03427 | 0.01234 | 2 |
| docstring | mean | k12_s811 / k12_s814 | 0.06245 | 0.2988 | 0.027 | 2 |
| acronym | mean | k2_s818 / k6_s812 | 0.02785 | 0.4786 | 0.1154 | 8.778 |
| Panel / pair | Shared | Median [Q1, Q3] (pp) | pp | Joint support |
| S3 Human suite R–C | 15 | 2.33 [0.50, 4.67] | 4 | 2/13 |
| S3 Human suite C–C | 597 | 3.67 [1.00, 18.67] | 271 | 152/472 |
| S3 InterpBench suite R–C | 60 | 2.19 [0.16, 7.15] | 25 | 38/60 |
| S3 InterpBench suite C–C | 1784 | 4.31 [0.74, 6.97] | 799 | 1158/1784 |
| S4 Human suite R–C | 22 | 1.00 [0.33, 4.83] | 6 | 5/16 |
| S4 Human suite C–C | 128 | 1.00 [0.67, 2.33] | 12 | 13/106 |