Transformer Interpretability

Latest papers 224

All topics
CardsList
  1. Beyond Linear Superposition: Discovering Climate Features in AI Weather Models with KAN-SAE

    May 17, 2026Minjong CheonTransformer InterpretabilityClimate Science

  2. How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning

    May 15, 2026Entang Wang, Yiwei Wang, Aleksandra Bakalova +1Transformer InterpretabilityTask Vectors

  3. On the Interpretability of Whisper Encodings Using Sparse Autoencoders

    May 12, 2026Dan Pluth, Zachary Nicholas Houghton, Yu Zhou +1Transformer InterpretabilityMechanistic Interpretability

  4. From Clever Hans to Scientific Discovery: Interpreting EEG Foundational Transformers with LRP

    May 12, 2026Justus Meyer zu Bexten, Nico Scherf, Bogdan Franczyk +1Transformer InterpretabilityEEG Decoding

  5. Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions

    May 11, 2026Andrew Lee, Fernanda Viégas, Martin WattenbergTransformer InterpretabilityNeural Representation Geometry

  6. From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models

    May 11, 2026Zehao Li, Yasuhiro Yoshikai, Shumpei Nemoto +2Transformer InterpretabilityLanguage Modeling

  7. Dissecting Jet-Tagger Through Mechanistic Interpretability

    May 11, 2026Saurabh Rai, Sanmay GangulyTransformer InterpretabilityAttention Head Analysis

  8. fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery

    May 10, 2026Andreas D. Demou, Panagiotis Koromilas, James Oldfield +2Transformer InterpretabilitySparse Autoencoders

  9. Attention Sinks in Diffusion Transformers: A Causal Analysis

    May 10, 2026Fangzheng Wu, Brian SummaTransformer InterpretabilityDiffusion Transformer

  10. Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression

    May 8, 2026Mingsong Yan, Dongyang Li, Charles Kulick +1Softmax AttentionTransformer Interpretability

  11. Belief or Circuitry? Causal Evidence for In-Context Graph Learning

    May 8, 2026Katharine Kowalyshyn, Timothy Duggan, Daniel Little +1Transformer InterpretabilityLLM Reasoning with Graphs

  12. Context-Gated Associative Retrieval: From Theory to Transformers

    May 8, 2026Moulik Choraria, Argyrios Gerogiannis, Vidhata Jayaraman +2Transformer InterpretabilityIn-Context Learning

  13. Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models

    May 7, 2026Amir Rezaei Balef, Mykhailo Koshil, Katharina EggenspergerTabular Foundation ModelsTransformer Interpretability

  14. From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features

    May 7, 2026Ruben Fernandez-Boullon, Pablo Magariños-Docampo, Javier Perez-RoblesTransformer InterpretabilitySparse Autoencoders

  15. The Metagame of Interpretability and Meta-Attributions

    May 7, 2026Hubert Baniecki, Przemyslaw Biecek, Fabian FumagalliTransformer InterpretabilityFeature Interaction Modeling

  16. Playing the network backward: A Game Theoretic Attribution Framework

    May 7, 2026Jakob Paul Zimmermann, Jim Berend, Georg Loho +2Gradient-Based AttributionTransformer Interpretability

  17. Metonymy in vision models undermines attention-based interpretability

    May 7, 2026Ananthu Aniraj, Cassio F. Dantas, Dino Ienco +2Transformer InterpretabilityDisentangled Representation Learning

  18. Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training

    May 7, 2026Hang Chen, Jiaying Zhu, Hongyang Chen +3Supervised Fine-TuningTransformer Interpretability

  19. Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers

    May 5, 2026Hao Yan, Haolin Yang, Yiqiao ZhongTransformer InterpretabilityNeural Representation Geometry

  20. Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers

    May 5, 2026Mathis Immertreu, Fitim Abdullahu, Thomas Kinfe +3Transformer InterpretabilityVision-Language Models

  21. Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models

    May 3, 2026Kotaro Furuya, Takahito TanimuraTransformer InterpretabilityLLM Evaluation