LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. From Directions to Regions: Decomposing Activations in Language Models via Local Geometry

    Feb 2, 2026Or Shafran, Shaked Ronen, Omri Fahn +3Language Model SteeringLLM Interpretability

  2. There Is More to Refusal in Large Language Models than a Single Direction

    Feb 2, 2026Faaiz Joad, Majd Hawasly, Sabri Boughorbel +2LLM InterpretabilityLLM Refusal Behavior

  3. Functional Subspace, where language models can use vector algebra to solve problems

    Feb 2, 2026Jung H. Lee, Sujith VijayanLLM InterpretabilityIn-Context Learning

  4. Tracing the Latent Threads: A Mechanistic Study of How LLMs Represent and Operationalize Race and Ethnicity Cues

    Jan 19, 2026Shiyue Hu, Ruizhe Li, Yanjun GaoLLM InterpretabilityMechanistic Interpretability

  5. To Copy or Not to Copy: Copying Is Easier to Induce Than Recall

    Jan 17, 2026Mehrdad Farahani, Franziska Penzkofer, Richard JohanssonKnowledge Conflicts in Language ModelsLLM Interpretability

  6. Relational Linearity is a Predictor of Hallucinations

    Jan 16, 2026Yuetian Lu, Yihong Liu, Sebastian Gerstner +3LLM InterpretabilityHallucination in Language Models

  7. Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

    Jan 16, 2026Lecheng Yan, Ruizhe Li, Guanhua Chen +5Reward HackingMemorization in Language Models

  8. Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models

    Jan 12, 2026Zhenghao He, Guangzhi Xiong, Bohan Liu +2LLM InterpretabilityCoT Reasoning

  9. Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

    Jan 6, 2026Harshvardhan Saini, Yiming Tang, Dianbo LiuLanguage Model SteeringLLM Interpretability

  10. Base Models Know How to Reason, Thinking Models Learn When

    Oct 8, 2025Constantin Venhoff, Iván Arcuschin, Philip Torr +2LLM InterpretabilityRL for Language Model Reasoning

  11. Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning

    Sep 1, 2025Kaviraj Pather, Elena Hadjigeorgiou, Arben Krasniqi +4LLM InterpretabilityHuman-in-the-Loop AI

  12. Unraveling the cognitive patterns of Large Language Models through module communities

    Aug 25, 2025Kushal Raj Bhandari, Pin-Yu Chen, Jianxi GaoComputational Cognitive ModelingLLM Interpretability

  13. Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

    Aug 22, 2025David Chanin, Adrià Garriga-AlonsoActivation SparsityLLM Interpretability

  14. BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection

    Aug 12, 2025Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3Social Bias in Language ModelsLLM Interpretability

  15. Too Categorical to be Human: Emotion Concepts in LLMs and Humans

    Aug 7, 2025Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig +4LLM InterpretabilityEmotion Representation in Language Models

  16. LLMs Encode Harmfulness and Refusal Separately

    Jul 16, 2025Jiachen Zhao, Jing Huang, Zhengxuan Wu +2LLM InterpretabilityLLM Refusal Behavior

  17. Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages

    Jul 15, 2025Lyzander Marciano Andrylie, Inaya Rahmanisa, Mahardika Krisna Ihsani +3Multilingual Language ModelsRepresentation Learning

  18. On the Effect of Uncertainty on Layer-wise Inference Dynamics

    Jul 9, 2025Sunwoo Kim, Haneul Yoo, Alice OhLLM InferenceLLM Interpretability

  19. The Trilemma of Truth in Large Language Models

    Jun 30, 2025Germans Savcisens, Tina Eliassi-RadLLM InterpretabilityLLM Reliability

  20. InverseScope: Scalable Activation Inversion for Interpreting Large Language Models

    Jun 9, 2025Yifan Luo, Zhennan Zhou, Bin DongRepresentation GeometryLLM Interpretability

  21. When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning

    May 21, 2025Rongzhi Zhu, Yi Liu, Jiancheng Wang +6Overthinking in Language ModelsLLM Interpretability

  22. Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations

    Apr 7, 2025Pedro Ferreira, Wilker Aziz, Ivan TitovCoT FaithfulnessReward Hacking

  23. Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

    Mar 17, 2025Noah Y. Siegel, Nicolas Heess, Maria Perez-Ortiz +1Faithfulness of Language Model ExplanationsLLM Interpretability

  24. LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

    Feb 20, 2025Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev +4Transformer InterpretabilityLong-Context Language Modeling

  25. Tokens, the oft-overlooked appetizer: Large language models, the distributional hypothesis, and meaning

    Dec 14, 2024Julia Witte Zimmerman, Denis Hudon, Kathryn Cramer +9LLM InterpretabilityDistributional Semantics

  26. Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

    Oct 3, 2024Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva +5Self-AttentionMultiple-Choice Question Answering

  27. How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism

    Date pendingElisabetta Rocchetti, Alfio FerraraLLM InterpretabilityInstruction Following