Mechanistic Interpretability

Recent momentum

-70%

6 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

2 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

6 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

154 papers

Latest in Mechanistic Interpretability

Open your feed →
CardsList
  1. Decoding Task Progress from VLA Representations

    Aug 13, 2026Atiksh Bhardwaj, Edward Weiyi Duan, Prithwish Dan +2Vision-Language-ActionVisuomotor Policy

  2. Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

    Jun 30, 2026Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo +5Mechanistic InterpretabilityAttribution Methods

  3. PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

    Jun 25, 2026Giosue Migliorini, Aristofanis Rontogiannis, Grigori Guitchounts +3ProteinSparse Autoencoders

  4. Effects of sparsity and superposition on loss in simple autoencoders

    Jun 16, 2026Mriganka Basu Roy Chowdhury, Eric McLaughlin WeinerSuperpositionAutoencoder

  5. Demystifying Variance in Circuit Discovery of LLMs

    Jun 15, 2026Frank Zhengqing Wu, Francesco Tonin, Volkan CevherMechanistic InterpretabilityCircuits