LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination

    Jun 30, 2026Vijay Vankadaru, Asha Matthews, Tanya Roosta +1LLM Hallucination MitigationLLM Interpretability

  2. C2^{2}R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders

    Jun 29, 2026Haoran Jin, Xiting Wang, Shijie Ren +2LLM InterpretabilityRepresentation Disentanglement

  3. Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression

    Jun 29, 2026Shuochen Chang, Qingyang Liu, Shaobo Wang +8LLM InterpretabilityEfficient Language Model Reasoning

  4. Mechanistically Eliciting Latent Behaviors in Language Models

    Jun 28, 2026Andrew Mack, Nina Panickssery, Alexander Matt TurnerLLM AlignmentLanguage Model Safety Evaluation

  5. The strength of clinical evidence is recoverable from language model representations but not from their stated grades

    Jun 27, 2026Soroosh Tayebi ArastehLLM InterpretabilityLanguage Model Calibration

  6. Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

    Jun 27, 2026David Courtis, Ting HuLanguage Model SteeringLLM Interpretability

  7. Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

    Jun 26, 2026Kevin Der, Harish Kamath, Ben ThompsonLLM InterpretabilitySparse Autoencoders

  8. VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

    Jun 26, 2026Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah +1LLM InterpretabilitySparse Autoencoders

  9. LMs as Task-Specific Knowledge Bases: An Interpretability Analysis

    Jun 25, 2026Amit Elhelo, Amir Globerson, Mor GevaLLM InterpretabilityFactual Knowledge in Language Models

  10. Forecasting With LLMs: Improved Generalization Through Feature Steering

    Jun 25, 2026Humzah Merchant, Bradford LevyLanguage Model SteeringLLM Interpretability

  11. Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

    Jun 25, 2026Sinie van der Ben, Raphaël Baur, Yannick Metz +1LLM InterpretabilityValence-Arousal Modeling

  12. Discovering Millions of Interpretable Features with Sparse Autoencoders

    Jun 25, 2026XinYang He, Wei Wang, Bing Zhao +5LLM InterpretabilitySparse Autoencoders

  13. Localizing RL-Induced Tool Use to a Single Crosscoder Feature

    Jun 25, 2026Andrii Shportko, Shubham Bhokare, Ahmed Zeyad A Alzahrani +3LLM InterpretabilityLLM Tool Use

  14. What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics

    Jun 23, 2026Sofiia Nikolenko, Michele Papucci, Mina Rezaei +1LLM InterpretabilityLLM Jailbreak Attacks

  15. Detecting and Controlling Sycophancy with Cascading Linear Features

    Jun 23, 2026Maty Bohacek, Rishub Jain, Nicholas Dufour +3LLM SycophancyLanguage Model Steering

  16. Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration

    Jun 23, 2026Pietro Tropeano, Maria Maistro, Tuukka Ruotsalo +1LLM PruningLLM Interpretability

  17. Evidence for feature-specific error correction in LLMs

    Jun 23, 2026Francisco Ferreira da Silva, Stefan HeimersheimLLM InterpretabilityLanguage Model Robustness

  18. Probing the Misaligned Thinking Process of Language Models

    Jun 23, 2026Kaiwen Zhou, Constantin Venhoff, Jonathan Michala +2LLM AlignmentLLM Auditing

  19. Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior

    Jun 22, 2026Christopher J. Anders, Henrique Da Silva Gameiro, Nico Daheim +1LLM AuditingLLM Interpretability

  20. Evaluation Awareness Is Not One Capability: Evidence from Open Language Models

    Jun 22, 2026Nilesh Nayan, Aishwarya Sampath Kumar, Rishiraj Girmal +5Language Model Safety EvaluationLLM Interpretability

  21. ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning Models

    Jun 22, 2026Jun Zhang, Jiasheng Zheng, Boxi Cao +5LLM AuditingLLM Interpretability

  22. Abstract representational geometry supports inference in large language models

    Jun 22, 2026Yunan Zeng, Yuwang WangLLM InferenceLLM Interpretability

  23. Exposing the Illusion of Erasure in Knowledge Editing for LLMs

    Jun 22, 2026Advik Raj Basani, Anshuman ChhabraKnowledge EditingKnowledge Conflicts in Language Models