Model Interpretability

Recent momentum

+38%

22 papers in the last 28 days · 0.4% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

9 new papers

A weekly snapshot of new work published in Model Interpretability.

Period ending 2026-09-14

8 new papers

A weekly snapshot of new work published in Model Interpretability.

Period ending 2026-09-07

5 new papers

A weekly snapshot of new work published in Model Interpretability.

269 papers

Latest in Model Interpretability

  1. When and How Long? The Readout-Mediator Angle in Temporal Reasoning

    May 27, 2026Shreyas Fadnavis, Praitayini Kanakaraj, Felix WyssReadoutModel Interpretability

  2. Learning to Translate from Soft to Hard LLM Prompts

    May 26, 2026Pitipat Kongsomjit, Suryansh Goyal, Jacob WhitehillMachine Translation ModelsPrompt Tuning

  3. Building Better Activation Oracles

    May 23, 2026Jan Bauer, Celeste De Schamphelaere, Adam Karvonen +2Model ActivationsModel Interpretability

  4. HiRes: Inspectable Precedent Memory for Reaction Condition Recommendation

    May 20, 2026Shreyas Vinaya Sathyanarayana, Raja Sekhar Pappala, Deepak WarrierSequential RecommendationModel Interpretability

  5. Where Does Authorship Signal Emerge in Encoder-Based Language Models?

    May 19, 2026Francis Kulumba, Guillaume Vimont, Laurent Romary +1AuthorshipEncoding Models

  6. iPOE: Interpretable Prompt Optimization via Explanations

    May 18, 2026Jiahui Li, Yarik Menchaca Resendiz, Sean Papay +1Prompt EngineeringPrompt Optimization

  7. Diversified Residual Symbolic Regression

    May 15, 2026Koki Ikeda, Masahiro Nomura, Ryoki HamanoSymbolic RegressionKernel Ridge Regression