Transformer Interpretability

Latest papers 224

All topics
CardsList
  1. Universal Redundancies in Time Series Foundation Models

    Feb 2, 2026Anthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai +1Transformer InterpretabilityKernel Regression

  2. Equivalence of Context and Parameter Updates in Modern Transformer Blocks

    Nov 22, 2025Adrian Goldwaser, Michael Munn, Javier Gonzalvo +1Transformer InterpretabilityTransformer FFNs

  3. Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task

    Oct 21, 2025Brady Bhalla, Honglu Fan, Nancy Chen +1Transformer InterpretabilityRepresentation Learning

  4. Deductive Logic in Language Models: Horizontal vs Vertical Reasoning

    Oct 10, 2025Davide Maltoni, Matteo FerraraTransformer InterpretabilityCoT Reasoning

  5. Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge

    Sep 29, 2025Pengxiao Lin, Zheng-An Chen, Zhi-Qin John XuTransformer InterpretabilityOOD Generalization

  6. Cross-Attention is Half Explanation in Speech-to-Text Models

    Sep 22, 2025Sara Papi, Dennis Fucci, Marco Gaido +2Transformer InterpretabilityCross-Attention

  7. REAL: Reading Out Transformer Activations for Precise Localization in Language Model Steering

    Jun 10, 2025Li-Ming Zhan, Bo Liu, Chengqiang Xie +2Transformer InterpretabilityLanguage Model Steering

  8. LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

    Feb 20, 2025Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev +4Transformer InterpretabilityLong-Context Language Modeling

  9. A Mechanistic Study of Transformers Training Dynamics

    Oct 31, 2024Ambroise Odonnat, Wassim Bouaziz, Vivien CabannesTransformer InterpretabilityAttention Mechanisms

  10. Tracing Computation Density in LLMs

    Date pendingCorentin Kervadec, Iuliia Lysova, Iuri Macocco +2Transformer InterpretabilityLLM Interpretability