Transformer Interpretability

Latest papers 224

All topics
CardsList
  1. How Token Influence Decays with Distance: A Green-Function View of Trained Language Models

    Jun 28, 2026Matthias Brändel, Stephan Köhler, Oliver RheinbachTransformer InterpretabilityAutoregressive Language Modeling

  2. From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection

    Jun 26, 2026Yasaman Haghbin, Sina Rashidi, Ali Zolnour +6Transformer InterpretabilityExplainable Artificial Intelligence

  3. At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

    Jun 24, 2026Praneet Suresh, Jack Stanley, Sonia Joseph +2Distribution Shift RobustnessTransformer Interpretability

  4. LIG: Layer-wise Integrated Gradients for Within-Layer Flow Analysis in Transformers

    Jun 19, 2026Eight Suzuki, Hideitsu Hino, Noboru MurataGradient-Based AttributionTransformer Interpretability

  5. Explaining Attention with Program Synthesis

    Jun 17, 2026Amiri Hayes, Belinda Z Li, Jacob AndreasTransformer InterpretabilityAttention Head Analysis

  6. Closing the Loop: PID Feedback Control for Interpretable Activation Steering in Symbolic Music Generation

    Jun 17, 2026Ioannis Prokopiou, Pantelis Vikatos, Maximos Kaliakatsos-Papakostas +2Transformer InterpretabilityControllable Music Generation

  7. When Does Depth Matter For In-Context Learning? Adaptive Inference in Deep Transformers

    Jun 15, 2026Ravin Raj, Gautam ReddyTransformer InterpretabilityAdaptive Inference

  8. Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

    Jun 12, 2026Ravi Ranjan, Utkarsh Grover, Xiaomin Lin +1Transformer InterpretabilityExplainable Artificial Intelligence

  9. Decompose Sparsely Where You Should, Absorb Densely Where You Should No

    Jun 12, 2026Ruixuan Deng, Zehao Jin, Zekun Wang +1Transformer InterpretabilityRepresentation Learning

  10. Explaining RhythmFormer: A Systematic XAI Analysis of Periodic Sparse Attention for Remote Photoplethysmography

    Jun 11, 2026Louis Chen, Torbjörn E. M. NordlingRemote PhotoplethysmographyTransformer Interpretability

  11. Layer-Resolved Optimal Transport for Hallucination Detection in NMT and Abstractive Summarization

    Jun 11, 2026Mariia Onyshchuk, Maksym-Vasyl Tarnavskyi, Marta SumykTransformer InterpretabilityHallucination Detection

  12. Vision Transformers for Face Recognition Need More Registers

    Jun 10, 2026Tahar Chettaoui, Guray Ozgur, Eduarda Caldeira +2Transformer InterpretabilityTransformer

  13. Integrating gene regulatory priors into Transformer attention with scTransformer for interpretable scRNA-seq analysis

    Jun 8, 2026Mikele Milia, Louis Fabrice Tshimanga, Henning Mueller +2Transformer InterpretabilityTransformer Attention

  14. Trajectory Geometry of Transformer Representations Across Layers

    Jun 8, 2026Vishal Pandey, Gopal Singh, Yacine MahdidTransformer InterpretabilityNeural Representation Geometry

  15. A Unifying View of Attention Sinks: From Mechanisms to Architectural Interventions

    Jun 6, 2026Lukas Fesser*, Mozes Jacobs*, Thomas Fel* +2Transformer InterpretabilityVision Transformer

  16. Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers

    Jun 4, 2026Tang Li, Yanlin Chen, Mengmeng Ma +1Transformer InterpretabilityVision Transformer

  17. Where does Absolute Position come from in decoder-only Transformers?

    Jun 4, 2026Valeria Ruscio, Umberto Nanni, Fabrizio SilvestriRotary Positional EmbeddingsTransformer Interpretability

  18. EIVE: End-to-End Instance-Specific Visual Explanations for Detection Transformers

    Jun 1, 2026Jianlin Xiang, Yanshan Li, Linhui DaiTransformer InterpretabilityFeature Attribution

  19. Assign and Add: A Mechanistic Study of Compositional Arithmetic

    May 29, 2026Brady Exoo, Alberto Bietti, John SousTransformer InterpretabilityTransformer