Transformer Interpretability

Latest papers 224

All topics
CardsList
  1. LAWFUL: Law-Aligned Witness for Faithful Use of Latents

    Jul 26, 2026Kevin Chen, Kenneth W. Parker, Anish AroraTransformer InterpretabilityMechanistic Interpretability

  2. Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

    Jul 25, 2026Dhruvil S, Fenil Sojitra, Ravirajsinh ChauhanTransformer InterpretabilityLow-Rank Attention

  3. Scaling Interpretable Transformers with Parity Bottleneck Layers

    Jul 22, 2026Andrew Mack, Kraig Yuheng Tou, Mark Henry +2Feature SuperpositionTransformer Interpretability

  4. Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

    Jul 21, 2026Maohua Li, Qirui Li, Yanke Zhou +10Transformer InterpretabilityDiffusion Transformer

  5. Circuit Claims Depend on What Is Extracted and How It Is Compared

    Jul 21, 2026Yang Sheng, Jie FuTransformer InterpretabilityCircuit Discovery

  6. For What Reason? Interpreting Models' Encoding of Causation and Antithesis

    Jul 20, 2026Abhidip Bhattacharyya, Shira WeinTransformer InterpretabilityCausal Reasoning in Language Models

  7. Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies

    Jul 19, 2026Luyu Qiu, Jianing Li, Hwanhee Kim +4Transformer InterpretabilityLanguage Modeling

  8. What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach

    Jul 19, 2026Afiq Abdillah Effiezal Aswadi, Haotong Ma, Susan WeiTransformer InterpretabilityLatent Variable Models

  9. Laguerre Geometry for Interpreting Large Language Models

    Jul 12, 2026Chunwei Ma, Russell WolfingerTransformer InterpretabilityLLM Interpretability

  10. The RG-Flow Transformer: Encoding Scale-Free Dynamics in Scarce EEG

    Jul 11, 2026Dibakar SigdelTransformer InterpretabilityElectroencephalography

  11. Gradient-Skipping Relevance Propagation for Efficient Explainability of Vision Transformers

    Jul 11, 2026Christopher Buratti, Michele Marchetti, Federica Parlapiano +3Transformer InterpretabilityVision Transformer

  12. What Pixels Are Enough? SEAMS: Sufficiency Saliency via MSE-Preservation Soft-Masks

    Jul 10, 2026Magdalena Trędowicz, Łukasz Struski, Arkadiusz Lewicki +4Transformer InterpretabilitySaliency Map Evaluation

  13. Training, Reading, and Editing Legible Transformers

    Jul 9, 2026Mark OskinTransformer InterpretabilityTransformer FFNs

  14. Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

    Jul 9, 2026Bendegúz Váradi, Zoltán KmettyTransformer InterpretabilitySparse Autoencoders

  15. Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

    Jul 8, 2026Pranav Sawant, Jakub KrejčíTransformer InterpretabilityMechanistic Interpretability

  16. Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits

    Jul 8, 2026Zitong Andrew Chen, Junaid Hasan, Akhil Srinivasan +2Transformer InterpretabilityTransformer Attention

  17. Faithfulness to Refusal: A Causal Audit of Neuron Selectors

    Jul 6, 2026Ananth Eswar, Pratinav Seth, Utsav Avaiya +1Transformer InterpretabilityFeature Attribution

  18. Legible-by-Construction: Attention and End-to-End Transformers

    Jul 5, 2026Mark OskinTransformer InterpretabilityTransformer Attention

  19. Individual Parameters in Weight-Sparse Transformers Appear Interpretable

    Jul 3, 2026Arnau Marin-Llobet, Stefan HeimersheimTransformer InterpretabilityTransformer

  20. Induction Heads Interpolate N-Grams

    Jul 2, 2026Francesco D'Angelo, Oguz Kaan Yuksel, Swathi Shree Narashiman +1Transformer InterpretabilityIn-Context Learning

  21. Representation as a Bottleneck for Mechanistic Interpretability: The Manifestation Unit Protocol

    Jun 30, 2026Hussein Chouman, Wataru Sasaki, Tomokazu Matsui +2Transformer InterpretabilityMechanistic Interpretability