LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Analysis and Explainability of LLMs Via Evolutionary Methods

    Apr 27, 2026Shannon K. Gallagher, Swati Rallapalli, Tyler Brooks +3LLM Interpretability

  2. Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer

    Apr 27, 2026Shun Shao, Binxu Wang, Shay B. Cohen +2LLM InterpretabilityMechanistic Interpretability

  3. Knowledge Vector of Logical Reasoning in Large Language Models

    Apr 26, 2026Zixuan Wang, Yuanyuan LeiLLM InterpretabilityRepresentation Engineering

  4. Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

    Apr 26, 2026John Winnicki, Abeynaya Gnanasekaran, Eric DarveKG ConstructionLLM Interpretability

  5. From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks

    Apr 25, 2026Nilanjana Das, Mathew Dawit, Aman Chadha +1LLM InterpretabilitySparse Autoencoders

  6. Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization

    Apr 24, 2026Weixu Zhang, Ye Yuan, Changjiang Han +7LLM AlignmentLanguage Model Steering

  7. How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals

    Apr 24, 2026Dharshan Kumaran, Viorica Patraucean, Simon Osindero +2LLM Self-CorrectionLanguage Model Error Detection

  8. Shared Lexical Task Representations Explain Behavioral Variability In LLMs

    Apr 23, 2026Zhuonan Yang, Jacob Xiaochen Li, Francisco Piedrahita Velez +5Prompt SensitivityLLM Prompting

  9. Slot Machines: How LLMs Keep Track of Multiple Entities

    Apr 22, 2026Paul C. Bogdan, Jack LindseyLLM InterpretabilityLanguage Model Probing

  10. TabSHAP

    Apr 22, 2026Aryan Chaudhary, Prateek Agarwal, Tejasvi AlladiFeature AttributionSHAP Feature Attribution

  11. Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs

    Apr 22, 2026Krishiv Agarwal, Ramneet Kaur, Colin Samplawski +6Language Model Safety EvaluationLLM Auditing

  12. LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures

    Apr 22, 2026Yuhang Wu, Qinyuan Liu, Qiuyang Zhao +1LLM InterpretabilityLanguage Model Robustness

  13. Surrogate modeling for interpreting black-box LLMs in medical predictions

    Apr 22, 2026Changho Han, Songsoo Kim, Dong Won Kim +4Surrogate ModelingLLM Interpretability

  14. R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling

    Apr 22, 2026Aijia Cheng, Kailong Wang, Ling Shi +1RL for Language ModelsLLM Interpretability

  15. Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders

    Apr 21, 2026Het Patel, Tiejin Chen, Hua Wei +2LLM InterpretabilityLLM Reliability

  16. How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning

    Apr 21, 2026Haoyang Chen, Yi Liu, Jianzhi Shao +3LLM InterpretabilityMathematical Reasoning

  17. Cell-Based Representation of Relational Binding in Language Models

    Apr 21, 2026Qin Dai, Benjamin Heinzerling, Kentaro InuiLLM InterpretabilityRepresentation Probing

  18. Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks

    Apr 20, 2026Md Rysul Kabir, Zoran TiganjLanguage Model Safety EvaluationLLM Interpretability

  19. Understanding the Prompt Sensitivity

    Apr 20, 2026Yang Liu, Chenhui ChuPrompt SensitivityLLM Prompting

  20. Reasoning Models Know What's Important, and Encode It in Their Activations

    Apr 20, 2026Yaniv Nikankin, Martin Tutek, Tomer Ashuach +2LLM InterpretabilityLLM Reasoning

  21. Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs

    Apr 20, 2026Charles Ye, Bo Yuan, Lee SharkeyLLM InterpretabilityNeural Network Interpretability

  22. Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks

    Apr 20, 2026Rongyuan Tan, Jue Zhang, Zhuozhao Li +3LLM EvaluationFeature Attribution