LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. On Temporal Binding in Large Audio Language Models

    Sep 28, 2026Paul Primus, Gerhard WidmerLLM InterpretabilityAudio Reasoning

  2. When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety

    Sep 28, 2026Tianyi Guan, Jianhui Chen, Liangming PanLanguage Model Safety EvaluationLLM Interpretability

  3. Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation

    Sep 27, 2026Haiyan Zhao, Zirui Hei, Wei Shi +3LLM Hallucination MitigationLLM Interpretability

  4. Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

    Sep 24, 2026Huseyin Cavus, Sebin Sabu, Joshua Spear +2LLM InterpretabilityHallucination in Language Models

  5. Grammatical "grandmother neurons" are rare in LLMs

    Sep 24, 2026Linyang He, Nima MesgaraniLLM Interpretability

  6. Stream Recursion Model (SRM)

    Sep 23, 2026Asael Sorensen, Charles Brock, David Chamberlain +3LLM InterpretabilityMechanistic Interpretability

  7. Math Reasoning in LLMs is Organized by Approach, Not Topic

    Sep 22, 2026Sajad Goudarzi, Samaneh Zamanifard, Moloud Nasiri +1Mathematical Reasoning BenchmarksLLM Interpretability

  8. Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

    Sep 22, 2026Xiaoyu Luo, Tao Ren, Wenrui Yu +3LLM InterpretabilityCoT Reasoning

  9. The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts

    Sep 21, 2026Manjiang Yu, Hongji Li, Zihan Wang +5LLM InterpretabilityLinear Representation Hypothesis

  10. Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models

    Sep 21, 2026Star S. D. Liu, Xiyu Ding, Robert B. Barrett +3Cost-Sensitive LearningLLM Interpretability

  11. From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness

    Sep 19, 2026Qianli Wang, Yilong Wang, Dennis Wei +5CoT FaithfulnessLLM Interpretability

  12. Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?

    Sep 19, 2026Weihan Li, Xinlei Chen, Yuhan Song +2Language Model SteeringLLM Interpretability

  13. Xeno-Interpretability: Investigating the Alien Minds of LLMs

    Sep 17, 2026F. Pierucci, M. Bracale Syrnikov, M. Prandi +3LLM InterpretabilityMechanistic Interpretability

  14. The Role of Fine-grained Harm Signals in LLM Safety

    Sep 16, 2026Soyeon Park, Seogyeong Jeong, Sunwoo Kim +1LLM InterpretabilityLLM Safety

  15. Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence

    Sep 16, 2026Sebastian Gerstner, Hilal AlQuabeh, Kentaro Inui +1Transformer InterpretabilityTransformer FFNs

  16. An Empirical Study of Counterfactual Self-Explanations in LLMs

    Sep 15, 2026Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis +2Counterfactual ExplanationsFaithfulness of Language Model Explanations

  17. TAME: Token Attribution and Masking for Emergent misalignment

    Sep 15, 2026Md Rayhanul Masud, Md Rizwan ParvezLLM InterpretabilityLLM Fine-Tuning

  18. Interpreting and Steering LLM Agents for Social Simulations

    Sep 14, 2026Jiayue Gaveal Fan, Arul Murugan, Shreyas Krishnan +1LLM InterpretabilityHuman Behavior Simulation

  19. Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs

    Sep 14, 2026JuHeon Ha, Byounghan Lee, Yunseo Choi +1Language Model SteeringLLM Interpretability

  20. The Misery of Mechanistic Interpretability: A Formal Perspective

    Sep 14, 2026Tobias Ladner, Matthias AlthoffNeural Network VerificationLLM Interpretability

  21. Artificial entrepreneurial cognition: Locating and causally steering an opportunity recognition dial inside large language models (LLMs)

    Sep 14, 2026Christian Fisch, Angela Altmeier, Martin Obschonka +2LLM Interpretability

  22. MAxBench: A Multinomial Concept Recovery Benchmark

    Sep 14, 2026Divya Appapogu, Freya Behrens, Yonatan Belinkov +1LLM InterpretabilityRepresentation Geometry in Language Models

  23. From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

    Sep 12, 2026Wenkang Wei, Yuan Fang, Renhe Jiang +2LLM InterpretabilityFactual Knowledge in Language Models