Model Interpretability Methods

Recent momentum

-57%

12 papers in the last 28 days · 0.3% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

8 new papers

A weekly snapshot of new work published in Model Interpretability Methods.

Period ending 2026-09-07

5 new papers

A weekly snapshot of new work published in Model Interpretability Methods.

262 papers

Latest in Model Interpretability Methods

Open your feed →
CardsList
  1. MAxBench: A Multinomial Concept Recovery Benchmark

    Sep 14, 2026Divya Appapogu, Freya Behrens, Yonatan Belinkov +1Model Interpretability Methods

  2. XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

    Sep 10, 2026Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein +1XaiExplainable Artificial Intelligence

  3. LLM Layers Immediately Correct Each Other

    Sep 7, 2026Arjun Patrawala, Jiahai Feng, Erik Jones +1Model ActivationsResidual Stream

  4. TabSOM: A tabular-to-image encoding method based on self-organizing maps

    Aug 13, 2026David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara +3Tabular LearningModel Interpretability Methods

  5. Scaling Inherently Interpretable Language Models

    Aug 6, 2026Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail +7Model Interpretability MethodsDiffusion Language Models

  6. Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders

    Jul 23, 2026Pablo Santiago Potes Velasco, María del Mar García Matabanchoy, Óscar Julián Pérez Ladino +3Large Language Model BiasStereotype Mitigation

  7. Neural Feature Governance: Extending Atom Prevalence

    Jul 23, 2026Idris Karel Seunda Ekwe, Patrick Tenga Shako, Ernest Parfait FokouéBayesian Neural NetworksModel Interpretability Methods

  8. AIMO Interpretability Challenge

    Jul 15, 2026Michal Štefánik, Philipp Mondorf, Andreas Waldis +11Olympiad-Level ProblemModel Interpretability Methods