Model Interpretability

Recent momentum

+38%

22 papers in the last 28 days · 0.4% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

9 new papers

A weekly snapshot of new work published in Model Interpretability.

Period ending 2026-09-14

8 new papers

A weekly snapshot of new work published in Model Interpretability.

Period ending 2026-09-07

5 new papers

A weekly snapshot of new work published in Model Interpretability.

269 papers

Latest in Model Interpretability

Open your feed →
CardsList
  1. Tracing Persona Vectors Through LLM Pretraining

    May 13, 2026Viktor Moskvoretskii, Dominik Glandorf, Jorge Medina Moreira +2PersonalityLarge Language Model Safety

  2. Metaphor Is Not All Attention Needs

    May 12, 2026Olga Sorokoletova, Francesco Giarrusso, Giacomo De Luca +6Large Language Model JailbreaksJailbreaks

  3. Interpretability Can Be Actionable

    May 11, 2026Hadas Orgad, Fazl Barez, Tal Haklay +9Model Interpretability

  4. ProDG: Prototypes for Data-Free Generative Post-Hoc Explainability

    May 9, 2026Piotr Borycki, Magdalena Trędowicz, Jacek Tabor +2XaiModel Interpretability

  5. Attributions All the Way Down? The Metagame of Interpretability

    May 7, 2026Hubert Baniecki, Przemyslaw Biecek, Fabian FumagalliModel InterpretabilityGame Theory

  6. TabSHAP

    Apr 22, 2026Aryan Chaudhary, Prateek Agarwal, Tejasvi AlladiTabular LearningModel Interpretability