Mechanistic Interpretability

Recent momentum

-8%

11 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

4 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-14

2 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

6 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

158 papers

Latest in Mechanistic Interpretability

  1. Demystifying Variance in Circuit Discovery of LLMs

    Jun 15, 2026Frank Zhengqing Wu, Francesco Tonin, Volkan CevherMechanistic InterpretabilityCircuits

  2. The Rate-Distortion-Polysemanticity Tradeoff in SAEs

    May 14, 2026Tommaso Mencattini, Francesco Montagna, Francesco LocatelloSemantic RepresentationsMechanistic Interpretability

  3. The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime

    May 11, 2026Phongsakon Mark Konrad, Tim Lukas Adam, Ane Cathrine Holst Merrild +4Artificial Intelligence GovernanceVerifier