Mechanistic Interpretability

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

12 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-14

10 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

13 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Inside this field

Focused directions

434 papers

Latest in Mechanistic Interpretability

  1. Metaphor Is Not All Attention Needs

    May 12, 2026Olga Sorokoletova, Francesco Giarrusso, Giacomo De Luca +6Large Language Model JailbreaksJailbreaks

  2. Interpretability Can Be Actionable

    May 11, 2026Hadas Orgad, Fazl Barez, Tal Haklay +9Model Interpretability

  3. The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime

    May 11, 2026Phongsakon Mark Konrad, Tim Lukas Adam, Ane Cathrine Holst Merrild +4Artificial Intelligence GovernanceVerifier

  4. ProDG: Prototypes for Data-Free Generative Post-Hoc Explainability

    May 9, 2026Piotr Borycki, Magdalena Trędowicz, Jacek Tabor +2XaiModel Interpretability

  5. Attributions All the Way Down? The Metagame of Interpretability

    May 7, 2026Hubert Baniecki, Przemyslaw Biecek, Fabian FumagalliModel InterpretabilityGame Theory