LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Fast & Faithful Function Vectors

    Jun 3, 2026Minh An Pham, Anton Segeler, Thomas Wiegand +4Task VectorsLLM Interpretability

  2. Do Large Language Models Have Emotions?

    Jun 3, 2026Amit Goldenberg, James J. GrossLLM EvaluationLLM Interpretability

  3. Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling

    Jun 2, 2026Yifan Wang, Jinyi Mu, Mayank Jobanputra +5Pairwise Preference LearningReward Modeling

  4. Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics

    Jun 2, 2026Mutsumi Sasaki, Go kamoda, Ryosuke Takahashi +4Numerical Reasoning in Language ModelsLLM Evaluation

  5. Framing Migration News with LLMs: Structured CoT as a Support for Human Interpretation

    Jun 2, 2026David Alonso del Barrio, Jing Wen, Daniel Gatica-PerezLLM InterpretabilityCoT Reasoning

  6. When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models

    Jun 2, 2026Ding Zhang, Runtao Zhou, Wenqing Zheng +3LLM InterpretabilityMechanistic Interpretability

  7. A Close Look At World Model Recovery In Supervised Fine-Tuned LLM Planners

    Jun 2, 2026Patrick Emami, Nan Qiang, Peter GrafSupervised Fine-TuningLLM Planning

  8. Decomposing how prompting steers behavior

    Jun 2, 2026Fan L. Cheng, Nikolaus KriegeskorteLLM PromptingLLM Interpretability

  9. Multi-component Causal Tracing in Large Language Models

    Jun 2, 2026Zirui Yan, Dennis Wei, Dmitriy A. Katz +2LLM InterpretabilityCausal Reasoning in Language Models

  10. On the Persistent Effects of Lexicality in Large Language Models

    Jun 1, 2026Hammad Rizwan, Muhammad Umair Haider, Nishant Subramani +3LLM InterpretabilityLanguage Modeling

  11. An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

    May 31, 2026Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama +1Reasoning EvaluationLLM Interpretability

  12. Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention

    May 31, 2026Shuochen Chang, Tong Bai, Xiaofeng Zhang +5LLM InterpretabilityCausal Interventions in Language Models

  13. The Case for Model Science: Verify, Explore, Steer, Refine

    May 31, 2026Przemyslaw Biecek, Luca Longo, Jianlong Zhou +3LLM EvaluationLLM Interpretability

  14. Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales

    May 31, 2026Pingjun Hong, Benjamin RothFeature AttributionLLM Interpretability

  15. MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models

    May 31, 2026Partha Pratim Saha, Samarth Raina, Mayur Parvatikar +4LLM InterpretabilityPreference Alignment

  16. Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

    May 30, 2026Hwiyeong Lee, Ingyu Bang, Uiji Hwang +2Representation LearningLLM Interpretability

  17. The Latin Substrate: How Language Models Represent and Mediate Script Choice

    May 29, 2026Daniil Gurgurov, Alan Saji, Katharina Trinley +2Multilingual Language ModelsLLM Interpretability

  18. Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines

    May 29, 2026Mikkel Godsk Jørgensen, Lars Kai HansenLanguage Model SteeringLLM Interpretability

  19. The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models

    May 29, 2026Liuyuan Wen, Xun Zhu, Lihao Huang +2Neural Representation GeometryLLM Interpretability

  20. Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits

    May 29, 2026Soorya Ram Shimgekar, Agam Goyal, Amruta Parulekar +6Prompt SensitivityLLM Interpretability