LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts

    May 24, 2026Xiangtian Ji, Yuxin Chen, Zhengzhou Cai +3Fine-TuningLLM Interpretability

  2. Building Better Activation Oracles

    May 23, 2026Jan Bauer, Celeste De Schamphelaere, Adam Karvonen +2LLM InterpretabilityMechanistic Interpretability

  3. Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

    May 22, 2026Jinghan Jia, Joe Benton, Eric EasleyCoT FaithfulnessLLM Interpretability

  4. Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning

    May 22, 2026Muhammad Usama, Dong Eui ChangRepresentational Similarity AnalysisLLM Interpretability

  5. Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

    May 21, 2026Dongxin Guo, Jikun Wu, Siu Ming YiuNeural ProcessesLLM Interpretability

  6. Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion

    May 21, 2026Meimingwei Li, Yuanhao Ding, Esteban Garces Arias +1LLM InterpretabilityLLM Fine-Tuning

  7. Relational Linear Properties in Language Models: An Empirical Investigation

    May 21, 2026Giovanni Valer, Luigi Gresele, Marco Bronzini +1LLM InterpretabilityFactual Knowledge in Language Models

  8. Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

    May 21, 2026Jewon Yeom, Jaewon Sok, Heejun Kim +3LLM InterpretabilityHallucination in Language Models

  9. Probabilistic Attribution For Large Language Models

    May 20, 2026Shilpika Shilpika, Carlo Graziani, Bethany Lusch +2LLM InterpretabilityLLM Uncertainty Estimation

  10. Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs

    May 19, 2026Edie Pearman, Sophia Osborne, Mira Kandlikar-Bloch +3Gender Bias in Language ModelsLLM Interpretability

  11. Language models struggle with compartmentalization

    May 19, 2026Thomas Vincent Howe, David WingateMultilingual Language ModelsLLM Interpretability

  12. Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

    May 19, 2026Somnath Banerjee, Pranav Jha, Rima Hazra +1Multilingual Language Model EvaluationFaithfulness of Language Model Explanations

  13. Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

    May 18, 2026Maciej Chrabąszcz, Aleksander Szymczyk, Marcin Sendera +2Language Model Safety EvaluationLLM Interpretability

  14. Probing for Representation Manifolds in Superposition

    May 18, 2026Alexander ModellFeature SuperpositionNeural Representation Geometry

  15. iPOE: Interpretable Prompt Optimization via Explanations

    May 18, 2026Jiahui Li, Yarik Menchaca Resendiz, Sean Papay +1LLM InterpretabilityLLM-Assisted Annotation

  16. Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models

    May 18, 2026Junyao Yang, Chen Qian, Kun Wang +4LLM InterpretabilityRL for Language Model Reasoning

  17. Artificial Aphasias in Lesioned Language Models

    May 15, 2026Nathan Roll, Jill Kries, Laura Gwilliams +1LLM InterpretabilityCausal Interventions in Language Models

  18. Judge Circuits Explain Format-Induced Inconsistency in LLM-as-a-Judge

    May 15, 2026Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10LLM-as-a-JudgeLLM Interpretability

  19. Reasoning Models Don't Just Think Longer, They Move Differently

    May 14, 2026Anders Gjølbye, Lars Kai Hansen, Sanmi KoyejoLLM InterpretabilityRepresentation Geometry in Language Models

  20. Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance

    May 14, 2026Mahdi Naser-Moghadasi, Faezeh GhaderiLLM EvaluationDecoder-Only Language Models

  21. Non-linear Interventions on Large Language Models

    May 14, 2026Sangwoo KimLLM InterpretabilityCausal Interventions in Language Models

  22. Exploring Geographic Relative Space in Large Language Models through Activation Patching

    May 14, 2026Stef De Sabbata, Rahul Baiju, Stefano Mizzaro +1LLM Interpretability