LLM Interpretability

LLM: Large Language Model

Latest papers 546

All topics
CardsList
  1. Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology

    May 14, 2026Jesseba Fernando, Grigori GuitchountsDynamical SystemsTransformer

  2. Polar probe linearly decodes semantic structures from LLMs

    May 13, 2026Pablo J. Diego-Simón, Pierre Orhan, Emmanuel Chemla +2Representation LearningLinear Probing

  3. Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity

    May 13, 2026Cristian Hinostroza, Rodrigo Toro Icarte, Christ Devia +4LLM PruningLLM Interpretability

  4. Probing Persona-Dependent Preferences in Language Models

    May 13, 2026Oscar Gilg, Pierre Beckmann, Daniel Paleka +1LLM AlignmentLLM Interpretability

  5. Tracing Persona Vectors Through LLM Pretraining

    May 13, 2026Viktor Moskvoretskii, Dominik Glandorf, Jorge Medina Moreira +2Language Model PretrainingLLM Interpretability

  6. When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

    May 13, 2026Vardhan Dongre, Joseph Hsieh, Viet Dac Lai +3Self-AttentionLLM Interpretability

  7. Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs

    May 12, 2026Jingzhou Jiang, Yi Yang, Kar Yan TamText EmbeddingsLLM Pruning

  8. All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs

    May 12, 2026Xi Chen, Mingyu Jin, Jingcheng Niu +7Circuit DiscoveryLLM Interpretability

  9. GKnow: Measuring the Entanglement of Gender Bias and Factual Gender

    May 12, 2026Leonor Veloso, Hinrich SchützeGender Bias in Language ModelsLLM Interpretability

  10. Targeted Neuron Modulation via Contrastive Pair Search

    May 12, 2026Sam Herring, Jake Naviasky, Karan MalhotraLanguage Model SteeringLLM Interpretability

  11. Do Language Models Encode Knowledge of Linguistic Constraint Violations?

    May 12, 2026Hardy, Sebastian PadóLLM InterpretabilityMechanistic Interpretability

  12. Domain Restriction via Multi SAE Layer Transitions

    May 12, 2026Elias Shaheen, Avi MendelsonLLM InterpretabilityOOD Detection

  13. Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

    May 12, 2026Yilong Wang, Qianli Wang, Bohao Chu +3Multilingual Language ModelsCounterfactual Explanations

  14. Temporal Preference Concepts and their Functions in a Large Language Model

    May 11, 2026Ian Rios-Sialer, Shantanu Darveshi, Shuai Jiang +4LLM InterpretabilityTemporal Reasoning in Language Models

  15. Instructions Shape Production of Language, not Processing

    May 11, 2026Andreas Waldis, Leshem Choshen, Yufang Hou +1LLM InterpretabilityInstruction Following

  16. SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing

    May 11, 2026Mingxu Zhang, Yuhan Li, Lujundong Li +3Molecular OptimizationLLM Interpretability

  17. Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs

    May 11, 2026Krishak Aneja, Manas Mittal, Anmol Goel +2LLM AlignmentLLM Interpretability

  18. SLASH the Sink: Sharpening Structural Attention Inside LLMs

    May 11, 2026Yiming Liu, Bin Lu, Xinbing Wang +2LLM Reasoning with GraphsLLM Interpretability

  19. Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations

    May 11, 2026Su-Hyeon Kim, Yo-Sub HanLLM InterpretabilityTransfer Learning