Mechanistic Interpretability

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-14

10 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

13 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Inside this field

Focused directions

434 papers

Latest in Mechanistic Interpretability

Open your feed →
CardsList
  1. MAxBench: A Multinomial Concept Recovery Benchmark

    Sep 14, 2026Divya Appapogu, Freya Behrens, Yonatan Belinkov +1Model Interpretability Methods

  2. XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

    Sep 10, 2026Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein +1XaiExplainable Artificial Intelligence

  3. LLM Layers Immediately Correct Each Other

    Sep 7, 2026Arjun Patrawala, Jiahai Feng, Erik Jones +1Model ActivationsResidual Stream

  4. TabSOM: A tabular-to-image encoding method based on self-organizing maps

    Aug 13, 2026David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara +3Tabular LearningModel Interpretability Methods

  5. Decoding Task Progress from VLA Representations

    Aug 13, 2026Atiksh Bhardwaj, Edward Weiyi Duan, Prithwish Dan +2Vision-Language-ActionVisuomotor Policy

  6. Scaling Inherently Interpretable Language Models

    Aug 6, 2026Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail +7Model Interpretability MethodsDiffusion Language Models

  7. Search Strategies for Optimal Classification and Regression Trees

    Jul 30, 2026Jacobus G. M. van der Linden, Mim van den Bos, Emir DemirovićDecision TreesRandom Forest