Mechanistic Interpretability

Recent momentum

-70%

6 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

2 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

6 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

154 papers

Latest in Mechanistic Interpretability

Open your feed →
CardsList
  1. ATLAS: Active Theory Learning for Automated Science

    Jun 10, 2026Noémi Éltető, Nathaniel D. Daw, Kimberly L. Stachenfeld +1Generalizable Time-Series Behavioral ModelingActive Learning

  2. Mechanisms of Object Localization in Vision-Language Models

    May 19, 2026Timothy Schaumlöffel, Martina G. Vilas, Gemma RoigTarget LocalizationLocalization

  3. The Rate-Distortion-Polysemanticity Tradeoff in SAEs

    May 14, 2026Tommaso Mencattini, Francesco Montagna, Francesco LocatelloSparse AutoencodersMechanistic Interpretability

  4. The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime

    May 11, 2026Phongsakon Mark Konrad, Tim Lukas Adam, Ane Cathrine Holst Merrild +4Artificial Intelligence DeploymentVerification