LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance

    Oct 8, 2026Zheng Huang, Sansheng Cao, Enpei Zhang +7LLM ReasoningLLM Interpretability

  2. SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

    Oct 7, 2026Miao Yu, Hao Huang, Lu Yuan +3LLM Safety AlignmentLLM Alignment

  3. How Do LLMs Change Predictions Under Negation?

    Oct 7, 2026Jongwook Yoon, Jongwon Lim, Sungjib Lim +2LLM ReliabilityLLM Evaluation

  4. The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models

    Oct 7, 2026Zhe Yu, Wenpeng Xing, Yunzhao Wei +4LLM InterpretabilityRetrieval-Augmented Generation

  5. U-Space: Uncovering When and Why Uncertainty Arises in Language Models

    Oct 6, 2026Tobias Braun, Nils Loose, Alexander Herzog +4LLM Uncertainty EstimationConfidence Estimation in Language Models

  6. Latent space bias directions in LLMs capture confidence, not fairness

    Oct 6, 2026Stephanie Buttigieg, Maeve Madigan, Parameswaran Kamalaruban +1Debiased MLLanguage Model Steering

  7. Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models

    Oct 6, 2026Md Abdul Kadir, Md Mohasin Hossain, Daniel SonntagAttention Head AnalysisLLM Interpretability

  8. Do LLMs Act on What They Know? From Partner Representations to Cooperative Actions

    Oct 6, 2026Yuhwan Jeong, Jinnyeong Yang, Kuk-Jin YoonMulti-Agent LLM SystemsLLM Interpretability

  9. COMPASS: Finding Where Reasoning Lives in Language Models

    Oct 5, 2026Pratyay Dutta, Kowshik Thopalli, Vivek NarayanaswamyEfficient Language Model ReasoningActivation Steering

  10. Identifying Introspection From the Inside

    Oct 5, 2026David I. Atkinson, Dillon Plunkett, David BauLLM InterpretabilityLanguage Model Introspection

  11. Before Agent Tells The Lie: Has Deception Already Been Represented?

    Oct 5, 2026Xinling Li, Dadi Guo, Qingyu Liu +5LLM InterpretabilityDeception in Language Models

  12. Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

    Oct 1, 2026Tido Specht, Elias Benedict Krey, Nils Neukirch +1LLM AlignmentLLM Interpretability

  13. Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models

    Oct 1, 2026Darpan Aswal, Céline HudelotIntegrated GradientsLLM Interpretability

  14. Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining

    Oct 1, 2026Shengye Tao, Yinzhu Cheng, Haihua XieLLM InterpretabilityCausal Interventions in Language Models

  15. Interpreting Reasoning of Large Language Models via Partial Information Decomposition

    Sep 30, 2026Barproda Halder, Qiuyi Zhang, Sanghamitra DuttaLLM InterpretabilityEfficient Language Model Reasoning

  16. Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting

    Sep 30, 2026Liang Twist Shan, Tianyu Hu, Hao Yan +1Numerical Reasoning in Language ModelsLong-Context Retrieval

  17. Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

    Sep 29, 2026Zhenting Huang, Bo Jiang, Junnan Liu +2Activation SparsityLLM Interpretability

  18. XU-RS: Explaining Credal Width in Random-Set Language Models

    Sep 29, 2026David Achara, Maryam Sultana, Alexander D. Rast +1LLM InterpretabilityLLM Uncertainty Estimation

  19. A Polyphonic Conception of AI Understanding

    Sep 28, 2026Matthieu Queloz, Pierre BeckmannLLM InterpretabilityLLM Reliability

  20. Causal and Interpretable Structures in LLM Compositional Tasks

    Sep 28, 2026Gurbir Arora, Toni J. B. Liu, Jiajun Bao +2LLM InterpretabilityCausal Interventions in Language Models

  21. Late Attention Layers Alone Can Copy Entity Tokens, but Not Without Attending to Their Context

    Sep 28, 2026Muyu He, Yuchen Liu, Ran Tao +1Memorization in Language ModelsLLM Interpretability

  22. Signatures of semantic search in the activations of large language models

    Sep 28, 2026Luke Leckie, Peter M. Todd, Jacob G. FosterExploration-Exploitation TradeoffLLM Interpretability

  23. Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery

    Sep 28, 2026Xu Wang, Yifan Yang, TingHao YU +1Representation LearningLLM Interpretability

  24. Imprint Reader: From Weight-Update Readout to Behavioral Intervention

    Sep 28, 2026Guanxu Chen, Qihao Lin, Jing ShaoLLM InterpretabilityNeural Network Interpretability

  25. A mechanistic study of language model introspection

    Sep 28, 2026Jiahong Zou, Xiangkun Sun, Lingkai Kong +1Attention Head AnalysisLLM Interpretability

  26. Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices

    Sep 28, 2026Lucio La Cava, Andrea TagarelliMoral Reasoning in Language ModelsLLM Auditing