LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering

    Apr 19, 2026Li Zheng, Xin Zhang, Shuyi He +5Language Model SteeringLLM Interpretability

  2. AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency

    Apr 17, 2026Max Henning Höth, Kristian Kersting, Björn Deiseroth +1CoT FaithfulnessFeature Attribution

  3. Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures

    Apr 17, 2026Yutong Gao, Qinglin Meng, Yuan Zhou +1Transformer InterpretabilityLLM Interpretability

  4. Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms

    Apr 17, 2026Tanja Baeumel, Josef van Genabith, Simon OstermannLLM InterpretabilityMechanistic Interpretability

  5. LLM Reasoning Is Latent, Not the Chain of Thought

    Apr 17, 2026Wenshuo WangLLM InterpretabilityCoT Reasoning

  6. LLM attribution analysis across different fine-tuning strategies and model scales for automated code compliance

    Apr 16, 2026Jack Wei Lun Shi, Minghao Dang, Wawan Solihin +1Feature AttributionLLM Interpretability

  7. Mechanistic Decoding of Cognitive Constructs in Large Language Models

    Apr 16, 2026Yitong Shou, Manhao GuanLLM InterpretabilityCausal Interventions in Language Models

  8. What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

    Apr 9, 2026Stephen Cheng, Sarah Wiegreffe, Dinesh ManochaLanguage Model SteeringLLM Interpretability

  9. What are They Thinking? Delineation, Probing, and Tracking of Concepts in LLMs

    Apr 7, 2026Mohamed Abdelwahab, Michelle Yu Collins, Sihan Chen +5Linear ProbingLLM Interpretability

  10. GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics

    Apr 3, 2026Yujing Wang, Yuanbang Liang, Yukun Lai +2LLM InterpretabilityLLM Uncertainty Estimation

  11. From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents

    Mar 27, 2026Trilok Padhi, Ramneet Kaur, Krishiv Agarwal +9LLM InterpretabilityRepresentation Probing

  12. INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation

    Mar 23, 2026Alexandra Kuleshova, Andrei Volodichev, Daria Kotova +1Hallucination DetectionLLM Interpretability

  13. Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation

    Mar 20, 2026Lasse Marten Jantsch, Dong-Jae Koh, Seonghyeon Lee +1Transformer InterpretabilityLLM Interpretability

  14. ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs

    Mar 19, 2026Abhinaba Basu, Pavan ChakrabortyExplanation EvaluationLLM Interpretability

  15. Causal Tracing of Audio-Text Fusion in Large Audio Language Models

    Mar 14, 2026Wei-Chih Chen, Chien-yu Huang, Hung-yi LeeLLM InterpretabilityAudio Understanding

  16. World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models

    Mar 4, 2026Elan BarenholtzLLM World ModelsWord Embeddings

  17. Step-Level Sparse Autoencoder for Reasoning Process Interpretation

    Mar 3, 2026Xuan Yang, Jiayu Liu, Yuhang Lai +3LLM InterpretabilitySparse Autoencoders

  18. The GRADIEND Python Package: An End-to-End System for Gradient-Based Feature Learning

    Feb 27, 2026Jonathan Drechsel, Steffen HerboldLLM InterpretabilityLanguage Modeling

  19. Probing for Knowledge Attribution in Large Language Models

    Feb 26, 2026Ivo Brink, Alexander Boer, Dennis UlmerLLM InterpretabilityHallucination in Language Models

  20. Compressed Sensing for Capability Localization in Large Language Models

    Feb 11, 2026Anna Bair, Yixuan Even Xu, Mingjie Sun +1Attention Head AnalysisLLM Interpretability

  21. Patches of Nonlinearity: Instruction Vectors in Large Language Models

    Feb 8, 2026Irina Bigoulaeva, Jonas Rohweder, Subhabrata Dutta +1Transformer InterpretabilityLLM Interpretability

  22. CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation

    Feb 4, 2026Zhao Tong, Chunlin Gong, Yiping Zhang +5Language Model Safety EvaluationLLM Interpretability

  23. A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

    Feb 2, 2026Harry Mayne, Justin Singh Kang, Dewi Gould +3Faithfulness of Language Model ExplanationsLLM Interpretability