Transformer Interpretability

Latest papers 224

All topics
CardsList
  1. Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

    May 29, 2026Harry Jake Cunningham, Nicola Muca CironeTransformer InterpretabilityFeature Attribution

  2. Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

    May 28, 2026Pierre-Antoine Lequeu, Camille Barboule, Benjamin PiwowarskiTransformer InterpretabilityDisentangled Representation Learning

  3. Improving Adversarial Robustness of Attribution via Implicit Regularization

    May 28, 2026Amir Mehrpanah, Matteo Gamba, Hossein AzizpourTransformer InterpretabilityNeural Network Robustness

  4. Unsupervised Semantic Segmentation Facilitates Model Understanding

    May 28, 2026Xiaoyan Yu, Lisa Mais, Jannik Franzen +4Transformer InterpretabilityVision Transformer

  5. The Attentional White Bear Effect in Transformer Language Models

    May 27, 2026Rebecca Ramnauth, Brian ScassellatiTransformer InterpretabilityLLM Alignment

  6. Integrated and Cross-Architecture Interpretation of LLM Reasoning

    May 27, 2026Leonardo Matthew Yauw, Wei-Bin Kou, Yujiu YangTransformer InterpretabilityLLM Interpretability

  7. ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions

    May 27, 2026Prathyush Poduval, Calvin Yeung, Neel Desai +1Transformer InterpretabilityActivation Sparsity

  8. MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability

    May 25, 2026Barsat KhadkaTransformer InterpretabilityCircuit Discovery

  9. Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

    May 25, 2026Haiyan Zhao, Zirui He, Guanchu Wang +3Transformer InterpretabilityLLM Interpretability

  10. Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability

    May 24, 2026David N. Olivieri, Antonio F. Pérez RodríguezTransformer InterpretabilityMechanistic Interpretability

  11. Reading Task Failure Off the Activations: A Sparse-Feature Audit of GPT-2 Small on Indirect Object Identification

    May 21, 2026Mahdi NasermoghadasiTransformer InterpretabilitySparse Autoencoders

  12. Represented Is Not Computed: A Causal Test of Candidate Algorithmic Intermediates in a Transformer

    May 21, 2026Ishita Darade, Sushrut ThoratNumerical Reasoning in Language ModelsTransformer Interpretability

  13. Towards Explainability of SLMs by investigating Token Level Activation

    May 21, 2026Sayantani Ghosh, Rajashik Datta, Amit Kumar Das +1Transformer InterpretabilityFeature Attribution

  14. Post-Hoc Understanding of Metaphor Processing in Decoder-Only Language Models via Conditional Scale Entropy

    May 20, 2026Lawhori Chakrabarti, Jennifer Johnson-Leung, Bert Baumgaertner +3Figurative Language UnderstandingTransformer Interpretability

  15. Chessformer: A Unified Architecture for Chess Modeling

    May 18, 2026Daniel Monroe, George Eilender, Philip Chalmers +2Transformer InterpretabilityGame-Playing Agents