Mechanistic Interpretability

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

12 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-14

10 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

13 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Inside this field

Focused directions

434 papers

Latest in Mechanistic Interpretability

  1. Caption Bottleneck Models

    Jul 1, 2026Seref Baris Cagliyan, Umut Ozdemir, Merve Tapli +1Concept Bottleneck ModelsModel Interpretability

  2. Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

    Jun 30, 2026Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo +5Mechanistic InterpretabilityClosure

  3. Partition-Guided Distance Saliency: Bridging Decision and Objective Spaces in Many-Objective Optimization

    Jun 29, 2026Cláudio Lúcio do Val Lopes, Flávio Vinícius Cruzeiro Martins, Elizabeth Fialho WannerMulti-Objective OptimizationSaliency

  4. PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

    Jun 25, 2026Giosue Migliorini, Aristofanis Rontogiannis, Grigori Guitchounts +3Protein Data BankProtein

  5. Beyond Shapley: Efficient Computation of Asymmetric Shapley Values

    Jun 23, 2026Ezequiel Companeetz, Santiago Cifuentes, Sergio AbriolaShapley ValueInterpretable Models

  6. Interpreting Latent CoT Reasoning as Dynamical Systems

    Jun 20, 2026Sabari Iyyappan Duraipandian, Shreya Sanjay Boyane, Manju Nagesh +3Efficient Latent ReasoningModel Interpretability

  7. How Transparent is DiffusionGemma?

    Jun 18, 2026Joshua Engels, Callum McDougall, Bilal Chughtai +11TransparencyInterpretable Models

  8. Comparing Linear Probes with Mahalanobis Cosine Similarity

    Jun 17, 2026Zhuofan Josh Ying, Peter Hase, Nikolaus KriegeskorteCosine SimilarityModel Interpretability

  9. Demystifying Variance in Circuit Discovery of LLMs

    Jun 15, 2026Frank Zhengqing Wu, Francesco Tonin, Volkan CevherMechanistic InterpretabilityCircuits