Mechanistic Interpretability

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

12 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-14

10 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

13 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Inside this field

Focused directions

434 papers

Latest in Mechanistic Interpretability

  1. Building Better Activation Oracles

    May 23, 2026Jan Bauer, Celeste De Schamphelaere, Adam Karvonen +2Model ActivationsModel Interpretability

  2. HiRes: Inspectable Precedent Memory for Reaction Condition Recommendation

    May 20, 2026Shreyas Vinaya Sathyanarayana, Raja Sekhar Pappala, Deepak WarrierSequential RecommendationModel Interpretability

  3. Where Does Authorship Signal Emerge in Encoder-Based Language Models?

    May 19, 2026Francis Kulumba, Guillaume Vimont, Laurent Romary +1AuthorshipEncoding Models

  4. iPOE: Interpretable Prompt Optimization via Explanations

    May 18, 2026Jiahui Li, Yarik Menchaca Resendiz, Sean Papay +1Prompt EngineeringPrompt Optimization

  5. Diversified Residual Symbolic Regression

    May 15, 2026Koki Ikeda, Masahiro Nomura, Ryoki HamanoSymbolic RegressionKernel Ridge Regression

  6. The Rate-Distortion-Polysemanticity Tradeoff in SAEs

    May 14, 2026Tommaso Mencattini, Francesco Montagna, Francesco LocatelloSemantic RepresentationsMechanistic Interpretability

  7. Tracing Persona Vectors Through LLM Pretraining

    May 13, 2026Viktor Moskvoretskii, Dominik Glandorf, Jorge Medina Moreira +2PersonalityLarge Language Model Safety