LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Rational Sparse Autoencoder

    Jun 12, 2026Naiyu Yin, Yue YuActivation SparsityLLM Interpretability

  2. A Low-Rank Subspace Analysis of LLM Interventions

    Jun 12, 2026Angira Sharma, Christian Schroeder de Witt, Philip Torr +2LLM InterpretabilityLLM Refusal Behavior

  3. When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs

    Jun 12, 2026Boris Marinov, Angira Sharma, Christian Schroeder de Witt +3Multilingual Language ModelsLLM Interpretability

  4. Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry

    Jun 11, 2026Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee +3LLM InterpretabilityRepresentation Geometry in Language Models

  5. Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning

    Jun 11, 2026Zach Studdiford, Gary LupyanCognitive ModelingLLM Interpretability

  6. Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models

    Jun 11, 2026Daniel Scalena, Sara Candussio, Luca Bortolussi +3CoT FaithfulnessLLM Interpretability

  7. Localizing Anchoring Pathways in Language Models

    Jun 11, 2026Hillary N. Owusu, Sarah Wiegreffe, Naomi H. FeldmanNumerical Reasoning in Language ModelsLLM Interpretability

  8. ICA Lens: Interpreting Language Models Without Training Another Dictionary

    Jun 10, 2026Sida Liu, Feijiang HanIndependent Component AnalysisLLM Interpretability

  9. Bergson: An Open Source Library for Data Attribution

    Jun 10, 2026Lucia Quirke, Louis Jaburi, David Johnston +6Gradient-Based AttributionLLM Interpretability

  10. Forecasting Future Behavior as a Learning Task

    Jun 9, 2026Mosh Levy, Yoav Goldberg, Asa Cooper SticklandLLM InterpretabilityLarge Reasoning Models

  11. MIRAGE: A Polarity-Flipping Encoding Subspace in LLM Agents

    Jun 9, 2026Pratibha Revankar, Kargi Chauhan, Jihye Kim +3LLM InterpretabilityAI Agent Security

  12. Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

    Jun 8, 2026Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1TTS SynthesisLanguage Model Steering

  13. The Neutral Mask: How Alignment Training Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model

    Jun 8, 2026Wendy K. TamLLM AlignmentLLM Interpretability

  14. PRISM: Recovering Instruction Sets from Language Model Activations

    Jun 8, 2026Gilad Gressel, Rahul Pankajakshan, Julia Diament +3Neural DecodingLLM Interpretability

  15. The Amplifying Mirror: Locating and Steering the Partisan Direction inside a Large Language Model

    Jun 7, 2026Wendy K. TamLLM InterpretabilityPolitical Bias in Language Models

  16. Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models

    Jun 7, 2026Lucrezia Laraspata, Giovanna Castellano, Gennaro VessioKnowledge Conflicts in Language ModelsLLM Interpretability

  17. Inside the LLM Word Factory

    Jun 7, 2026Benzi Busigin, Yuval PinterLanguage Model DecodingLLM Interpretability

  18. Cross-LLM Consistency in Inference: Evidence from Shared Interactions

    Jun 6, 2026Siyu Lou, Yao Yan, Yuntian Chen +1LLM InferenceLLM Interpretability

  19. Beyond Post-hoc Explanation: Toward Glassbox AI via Probabilistic Mediation

    Jun 5, 2026Manuele LeonelliExplainable Artificial IntelligenceAI Accountability

  20. When Attribution Patching Lies: Diagnosis and a Second-Order Correction

    Jun 5, 2026Luyang Zhang, Jialu WangGradient-Based AttributionLLM Interpretability

  21. Interpreting Brain Responses to Language with Sparse Features from Language Models

    Jun 5, 2026Michael A. Lepori, Kendrick Kay, Greta TuckuteLLM InterpretabilitySparse Autoencoders

  22. LLM Self-Recognition: Steering and Retrieving Activation Signatures

    Jun 4, 2026Thibaud Ardoin, Jonas Schäfer, Gerhard WunderAI-Generated Text DetectionLLM Interpretability

  23. The Tell-Tale Norm: ℓ2\ell_2 Magnitude as a Signal for Reasoning Dynamics in Large Language Models

    Jun 4, 2026Jinyang Zhang, Hongxin Ding, Yue Fang +4LLM InterpretabilityTest-Time Scaling

  24. LLM Explainability with Counterfactual Chains and Causal Graphs

    Jun 4, 2026Nirit Nussbaum-Hoffer, Nitay Calderon, Liat Ein-Dor +1LLM Reasoning with GraphsLLM Interpretability