cs.LGSep 28, 2026

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

Authors: Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang, Luke Zhang, Haolin Ye, Xujie Si

Organizations: University of Toronto · McGill University · Tsinghua University

Abstract

Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

    May 8, 2026Michael Li, Nishant SubramaniMechanistic InterpretabilityCircuits

  2. Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

    Oct 1, 2026Chuqin Geng, Li Zhang, Haolin Ye +3Mechanistic InterpretabilityModel Discovery

  3. Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

    Jun 23, 2026Ayan Antik Khan, Harsh Kohli, Yuekun Yao +2Mechanistic InterpretabilityExplainable AI Methods