LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

    May 7, 2026Fahd Seddik, Fatemeh FardLLM EvaluationLLM Interpretability

  2. Gyan: An Explainable Neuro-Symbolic Language Model

    May 6, 2026Venkat Srinivasan, Vishaal Jatav, Anushka Chandrababu +1LLM InterpretabilityKnowledge Representation

  3. A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions

    May 5, 2026Qingzhuo Wang, Ruiyang Qin, Zhenxin Qin +2Feature Interaction ModelingLLM Interpretability

  4. Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

    May 4, 2026Yash Aggarwal, Atmika Gorti, Vinija Jain +3Moral Reasoning in Language ModelsSocial Bias in Language Models

  5. Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation

    May 4, 2026Francesco Sovrano, Gabriele Dominici, Marc LangheinrichLLM InterpretabilityMechanistic Interpretability

  6. How Language Models Process Negation

    May 4, 2026Zhejian Zhou, Tianyi Zhou, Robin Jia +1LLM InterpretabilityLanguage Modeling

  7. Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models

    May 3, 2026Kotaro Furuya, Takahito TanimuraTransformer InterpretabilityLLM Evaluation

  8. The Cylindrical Representation Hypothesis for Language Model Steering

    May 3, 2026Lang Gao, Jinghui Zhang, Wei Liu +7Language Model SteeringLLM Interpretability

  9. Reasoning emerges from constrained inference manifolds in large language models

    May 2, 2026Yanbiao Ma, Fei Luo, Linfeng Zhang +10LLM InferenceLLM Interpretability

  10. Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts

    May 1, 2026Sheridan Feucht, Tal Haklay, Usha Bhalla +9Numerical Reasoning in Language ModelsLLM Interpretability

  11. Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions

    Apr 30, 2026Jan Sobotka, Mustafa O. Karabag, Ufuk TopcuImperfect-Information GamesLLM Interpretability

  12. DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

    Apr 30, 2026Lifan Zheng, Xue Yang, Jiawei Chen +6LLM InterpretabilityPersonality Modeling in Language Models

  13. Modeling Clinical Concern Trajectories in Language Model Agents

    Apr 30, 2026Sukesh Subaharan, Venkatesan VS, Murugadasan P +3LLM InterpretabilityAI Agent Monitoring

  14. TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models

    Apr 30, 2026Amirreza Esmaeili, Fatemeh FardLLM InterpretabilityCode Generation

  15. Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models

    Apr 29, 2026Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter +3LLM InterpretabilityInstruction Following

  16. Semantic Structure of Feature Space in Large Language Models

    Apr 29, 2026Austin C. Kozlowski, Andrei BoutylineNeural Representation GeometryLLM Interpretability

  17. What Suppresses Nash Equilibrium Play in Large Language Models? Mechanistic Evidence and Causal Control

    Apr 29, 2026Paraskevas V. Lekeas, Giorgos StamatopoulosNash EquilibriumGame Theory

  18. MoRFI: Monotonic Sparse Autoencoder Feature Identification

    Apr 29, 2026Dimitris Dimakopoulos, Shay B. Cohen, Ioannis KonstasLLM Hallucination MitigationLLM Interpretability