LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Neuron-Level Interventions for Gendered and Gender-Neutral Generation in Language Models

    May 29, 2026Zhiwen You, Nafiseh Nikeghbal, Jana DiesnerGender Bias in Language ModelsLLM Interpretability

  2. Human-Alignment, Calibration, and Activation Patterns in Large Language Model Uncertainty

    May 29, 2026Kyle Moore, Jesse Roberts, Daryl Watson +2Confidence Estimation in Language ModelsLLM Interpretability

  3. Do Language Models Track Entities Across State Changes?

    May 28, 2026Zilu Tang, Qiao Zhao, Gabriel Franco +4LLM InterpretabilityState Tracking

  4. Latent Performance Profiling of Large Language Models

    May 28, 2026Tanmoy Chakraborty, Ayan Sengupta, Suparna Bhattacharya +7LLM EvaluationLLM Interpretability

  5. Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection

    May 28, 2026Syafiq Al Atiiq, Chun Zhou, Christian GehrmannLLM InterpretabilitySoftware Vulnerability Detection

  6. Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

    May 28, 2026David Fraile Navarro, Berardino Como, Jialei Sheng +2HealthcareClinical Triage

  7. When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception

    May 28, 2026Vahideh ZolfaghariLinear ProbingLanguage Model Safety Evaluation

  8. Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text

    May 27, 2026Tianyang Zhou, Wenbo Chen, Pierre Jinghong Liang +1LLM InterpretabilityText Classification

  9. Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing

    May 27, 2026Li Lei, Madalina Ciobanu, Qingqing Mao +1Task VectorsLLM Interpretability

  10. Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

    May 27, 2026Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta +4LLM InterpretabilityAdversarial Attacks on LLMs

  11. Cultural Binding Heads in Language Models

    May 27, 2026Avrile Floro, Luca BenedettoLLM InterpretabilityCultural Bias in Language Models

  12. Integrated and Cross-Architecture Interpretation of LLM Reasoning

    May 27, 2026Leonardo Matthew Yauw, Wei-Bin Kou, Yujiu YangTransformer InterpretabilityLLM Interpretability

  13. Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations

    May 27, 2026Simardeep Singh, Paras ChopraNeural Representation GeometryLLM Interpretability

  14. CAREF: Calibration-Aware Regularization for Explanation Faithfulness Without Rationale Supervision

    May 27, 2026Naphat Nithisopa, Teerapong PanboonyuenFaithfulness of Language Model ExplanationsLLM Interpretability

  15. Revealing Algorithmic Deductive Circuits for Logical Reasoning

    May 27, 2026Phuong Minh Nguyen, Tien Huu Dang, Naoya InoueAttention Head AnalysisLLM Interpretability

  16. Learning to Translate from Soft to Hard LLM Prompts

    May 26, 2026Pitipat Kongsomjit, Suryansh Goyal, Jacob WhitehillLLM PromptingLLM Interpretability

  17. The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context

    May 26, 2026Zhe Yu, Wenpeng Xing, Yunzhao Wei +4Memorization in Language ModelsLLM Interpretability

  18. Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations

    May 25, 2026Shanghao Li, Jinda Han, Yibo Wang +5LLM GroundingLLM Interpretability

  19. Can LLMs Introspect? A Reality Check

    May 25, 2026Shashwat Singh, Tal Linzen, Shauli RavfogelLLM EvaluationLLM Interpretability

  20. Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

    May 25, 2026Federico Torrielli, Peter Schneider-Kamp, Lukas Galke PoechConfidence Estimation in Language ModelsLLM Interpretability

  21. Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

    May 25, 2026Haiyan Zhao, Zirui He, Guanchu Wang +3Transformer InterpretabilityLLM Interpretability