cs.LGSep 14, 2026

The Misery of Mechanistic Interpretability: A Formal Perspective

Authors: Tobias LadnerMatthias Althoff

Organizations: Technical University of Munich, Germany

Abstract

Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inherently black boxes. To gain insights into these models, interpretable replacement networks (IRNs) are trained at all layers, exposing interpretable features through sparsely activated neurons. However, the faithfulness of an IRN is usually evaluated only empirically on clean data, and we show that even semantically minor input perturbations flip the dominant IRN features-and thus the human-understandable interpretation-across five open-weight model families (GPT-2 small, Gemma 2 2B, Gemma 3 1B, Llama 3.2 1B, R1-Distill-Qwen 1.5B). We propose the first formal verification framework for the faithfulness of an IRN, where reachability analysis certifies a sound upper bound of the faithfulness gap in adversarial scenarios. Moreover, we show that verification-aware training of IRNs substantially tightens this certified bound, restoring a feature-level interpretation that safety auditors can act on. Together, these results give, to the best of our knowledge, the first formal guarantees for mechanistic interpretability of large language models.

Explore similar work

CardsList
  1. Scaling Inherently Interpretable Language Models

    Aug 6, 2026Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail +7Model InterpretabilityTraining Language Models

  2. Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

    Jun 30, 2026Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo +5Mechanistic InterpretabilityClosure