Mechanistic Interpretability

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

12 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-14

10 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Period ending 2026-09-07

13 new papers

A weekly snapshot of new work published in Mechanistic Interpretability.

Inside this field

Focused directions

434 papers

Latest in Mechanistic Interpretability

  1. Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders

    Jul 23, 2026Pablo Santiago Potes Velasco, María del Mar García Matabanchoy, Óscar Julián Pérez Ladino +3Stereotype MitigationModel Interpretability

  2. Neural Feature Governance: Extending Atom Prevalence

    Jul 23, 2026Idris Karel Seunda Ekwe, Patrick Tenga Shako, Ernest Parfait FokouéBayesian Neural NetworksNeural Network

  3. Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy Library

    Jul 22, 2026Cayan Deniz Kucuktopana, Javier Fumanal-Idocin, Richard Pitts +1Interpretable ModelsExtended Version

  4. AIMO Interpretability Challenge

    Jul 15, 2026Michal Štefánik, Philipp Mondorf, Andreas Waldis +11Mathematical ReasoningOlympiad-Level Problem

  5. Noisy-Channel Minimum Bayes Risk Decoding

    Jul 6, 2026Yusuke Sakai, Hidetaka Kamigaito, Taro WatanabeChannel ModelingBayesian