Mechanistic Interpretability

Recent momentum

-70%

6 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

2 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

6 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

154 papers

Latest in Mechanistic Interpretability

Open your feed →
CardsList
  1. From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features

    May 7, 2026Ruben Fernandez-Boullon, Pablo Magariños-Docampo, Javier Perez-RoblesSparse AutoencodersCo-Occurrence

  2. There Will Be a Scientific Theory of Deep Learning

    Apr 23, 2026Jamie Simon, Daniel Kunin, Alexander Atanasov +11Deep Neural NetworkSingular Learning Theory

  3. Interpretability and Generalization Bounds for Learning Spatial Physics

    Jun 18, 2025Alejandro Francisco Queiruga, Theo Gutman-Solo, Shuai JiangGeneralization BoundsPhysics