LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. From Plausible to Actionable: A Position on LLM Self-Explanations

    Jul 17, 2026Elize Herrewijnen, Benedetta Muscato, Gizem Gezici +1CoT FaithfulnessFaithfulness of Language Model Explanations

  2. AIMO Interpretability Challenge

    Jul 15, 2026Michal Štefánik, Philipp Mondorf, Andreas Waldis +11Mathematical Reasoning BenchmarksLLM Interpretability

  3. Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

    Jul 13, 2026Zixiang Xu, Sixian Li, Huaxing Liu +4LLM-as-a-JudgeLLM Interpretability

  4. Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs

    Jul 12, 2026Shrestha Datta, Hongfu Liu, Anshuman ChhabraGradient-Based AttributionLLM Interpretability

  5. Laguerre Geometry for Interpreting Large Language Models

    Jul 12, 2026Chunwei Ma, Russell WolfingerTransformer InterpretabilityLLM Interpretability

  6. SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models

    Jul 11, 2026Dongxu Zhang, Yiding Sun, Zihao Guo +5LLM InterpretabilityActivation Steering

  7. Belief-reality separation lives in routing over a shared value slot in language models

    Jul 11, 2026Oliver Steele, Jiangtao Wen, Yuxing HanTheory of MindLLM Interpretability

  8. TypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code Models

    Jul 9, 2026Giuliano Gorgone, Fausto CarcassiLLM InterpretabilityRepresentation Geometry in Language Models

  9. Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

    Jul 8, 2026Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang +1LLM InterpretabilityAdversarial Attacks on LLMs

  10. Distributed Sparse Interventions in Language Models

    Jul 8, 2026Maximilian S. Ernst, Lorenz Linhardt, Aaron Peikert +1Language Model SteeringLLM Interpretability

  11. Dissociating the Internal Representations of Sycophancy in LLMs

    Jul 8, 2026Anthony Baez, Sheer Karny, Pat PataranutapornLLM SycophancyLLM Interpretability

  12. Reward Valuation in Large Language Models: Causal Induction of Anhedonia

    Jul 7, 2026Melika Honarmand, Samin Mahdipour Aghabagher, Martin SchrimpfReward ModelingCognitive Modeling

  13. How Much is Left? LLMs Linearly Encode Their Remaining Output Length

    Jul 6, 2026Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi +2LLM Interpretability

  14. Latent Programming Horizons in Coding Agents

    Jul 6, 2026André Silva, Han Tu, Martin MonperrusLLM InterpretabilityCoding Agents

  15. Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

    Jul 4, 2026Kareem Elozeiri, Mervat Abassy, Omar Kallas +4Arabic NLPLLM Interpretability

  16. Reading Between the Dots: Decoding Hidden Computation across Filler Tokens

    Jul 3, 2026Kaley Brauer, Claudio Mayrink Verdun, Samuel MarksLLM AuditingLLM Interpretability

  17. Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

    Jul 2, 2026Thomas WinningerLLM InterpretabilityLLM Safety

  18. Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates

    Jul 1, 2026Elias Najarro, Ane Espeseth, Eleni Nisioti +2Artificial LifeLLM Interpretability

  19. Understanding Large Language Models

    Jul 1, 2026Yannik Keller, Thomas EisenmannLLM Interpretability

  20. Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

    Jul 1, 2026Aryo Pradipta Gema, Beatrice Alex, Pasquale MinerviniLong-Context RetrievalAttention Head Analysis

  21. Shapley in Context: Explaining Financial Language with Domain Expertise

    Jul 1, 2026Dangxing Chen, Pengzhan GuoLLM InterpretabilityShapley Value Attribution

  22. Prototype Language Models

    Jul 1, 2026Dan Ley, Giang Nguyen, Himabindu Lakkaraju +1LLM InterpretabilityLanguage Modeling

  23. A Mechanistic View of Authority Hierarchy in LLM Sycophancy

    Jul 1, 2026Emil Joswin, Srujananjali Medicherla, Priyanka Mary MammenLLM SycophancyLLM Interpretability

  24. NeuroCogMap Reveals Cognitive Organization of Large Language Models

    Jul 1, 2026Zhongxiang Sun, Haolang Lu, Qiang Ma +11Cognitive ModelingLLM Interpretability