Mechanistic Interpretability

Momentum

20 papers in the last four weeks, up 54% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 175

All topics
CardsList
  1. Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

    Oct 1, 2026Chuqin Geng, Li Zhang, Haolin Ye +3Mechanistic InterpretabilityModel Discovery

  2. Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

    Oct 1, 2026Tido Specht, Elias Benedict Krey, Nils Neukirch +1Large Language Model AlignmentMultimodal Large Language Models

  3. When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task

    Sep 29, 2026Sai Sumedh R. Hindupur, Hadas Orgad, Thomas Fel +1Multilayer PerceptronsMechanistic Interpretability

  4. From Retrieval to Reasoning: Agentic Mechanism Prediction from Cell Painting Profiles

    Sep 29, 2026Jiayuan Chen, Botao Yu, Tianyu Liu +3PhenotypesCells

  5. Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

    Sep 28, 2026Li Zhang, Chuqin Geng, Mark Zhang +4Mechanistic InterpretabilityCircuits

  6. Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

    Sep 28, 2026Xu Wang, Difan Zou, Xuansheng WuSycophancyRefusals

  7. From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features

    Sep 28, 2026Dewen Liu, Zixuan Li, Jonathan Pan +4Mechanistic InterpretabilityModel Activations

  8. Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders

    Sep 28, 2026Zichao Yu, Qianshuo Ye, Xu Wang +1Efficient On-Policy DistillationOnline-Policy Distillation

  9. On Temporal Binding in Large Audio Language Models

    Sep 28, 2026Paul Primus, Gerhard WidmerLarge Audio Language ModelsMechanistic Interpretability

  10. Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

    Sep 24, 2026Huseyin Cavus, Sebin Sabu, Joshua Spear +2Generative HallucinationNeurons

  11. Stream Recursion Model (SRM)

    Sep 23, 2026Asael Sorensen, Charles Brock, David Chamberlain +3Mechanistic InterpretabilityRecursive Models

  12. Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders

    Sep 21, 2026Rithin Nagaraj, Rupa Laalasa Oruganti, Prerna Subhashchandra Kunder +1Transformer ArchitecturesTransformer Attention

  13. Topographic Training Concentrates Causal Circuits Without Improving Neuron Monosemanticity

    Sep 21, 2026Gautam Ranka, Shubham Santosh Pandere, Aiden DsouzaMechanistic InterpretabilityNeurons

  14. Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs

    Sep 20, 2026Edward G. Friedman, Xiangchen SongCircuitsMechanistic Interpretability

  15. Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

    Sep 17, 2026Noah Mamié, Laurin van den BerghProfileTransparency

  16. The Misery of Mechanistic Interpretability: A Formal Perspective

    Sep 14, 2026Tobias Ladner, Matthias AlthoffMechanistic InterpretabilityInterpretability

  17. Artificial entrepreneurial cognition: Locating and causally steering an opportunity recognition dial inside large language models (LLMs)

    Sep 14, 2026Christian Fisch, Angela Altmeier, Martin Obschonka +2Mechanistic InterpretabilityDial

  18. What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

    Sep 14, 2026Lingheng Du, Yiming Tang, Xufeng Duan +1Large Language Model Reinforcement LearningImproving Sparse Autoencoders

  19. On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing

    Sep 7, 2026Yuval Koren, Assaf Ben-Kish, Raja Giryes +2Dense Associative MemoryFactual Recall

  20. Large Language Models in Resolving Contextual Knowledge Conflicts

    Sep 2, 2026Xinye Yang, Zhenyang Liu, Ruisi Li +1LLM Reasoning StrategiesReasoning Skills

  21. S^3martCirc: Self-supervised Smart Circuit Discovery

    Sep 1, 2026Wendy Zheng, Yinhan He, Liang Wu +1Mechanistic InterpretabilityCircuits

  22. MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

    Aug 31, 2026Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang +5InterpretabilityMechanistic Interpretability

  23. The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal

    Aug 31, 2026Md Mokarram Chowdhury, Ernie Chang, Yang LiJailbreaksJailbreak Attacks

  24. RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons

    Aug 25, 2026Runyu Wang, Bo Liu, Xiaxin Zhang +6NeuronsMechanistic Interpretability

  25. Decoding Task Progress from VLA Representations

    Aug 13, 2026Atiksh Bhardwaj, Edward Weiyi Duan, Prithwish Dan +2Visuomotor PolicyMechanistic Interpretability

  26. Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

    Aug 13, 2026Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8DissociationLinguistics

  27. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

    Aug 12, 2026Mengru Wang, Junfeng Fang, Shuofei Qiao +16Artificial Intelligence SafetyArtificial Intelligence Scientists

  28. Measuring Semantic Abstractness of SAE Features via Nonlocality

    Aug 11, 2026Chuqiao Lin, Shivaji Sondhi, Xiao-Liang QiSparse Autoencoder FeaturesMechanistic Interpretability

  29. Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

    Aug 10, 2026Shuyi Miao, Wangjie Qiu, Pengyang Shao +4Multilingual SafetyLarge Language Model Safety

  30. MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering

    Aug 6, 2026Jakub Poćwiardowski, Mateusz ModrzejewskiText-To-MusicText Generation

  31. Sparse Weight Decomposition for Efficient Circuit Extraction

    Aug 4, 2026Chuanhao Yan, Xuhan Huang, Yawen Duan +4Transformer ArchitecturesModel Weights

  32. DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning

    Aug 1, 2026An Lanji, Dawei Liu, Jin Li +3Visual ReasoningRecent Vision-Language Models

  33. Emergent Latent-State Computation under Stochastic Volatility

    Jul 28, 2026Xiaoyu Huang, Lulu WangLatent StatesIntermediate Latent States

  34. Do LLMs Know Their Vulnerable Scenarios?

    Jul 26, 2026Ziheng Peng, Huiqi Deng, Haoran Jing +5Mechanistic Interpretability

  35. Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation

    Jul 25, 2026Gonzalo A. RuzRegulatory NetworksSurrogate Models

  36. Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

    Jul 25, 2026Dhruvil S, Fenil Sojitra, Ravirajsinh ChauhanMulti-Head AttentionDeepseek

  37. From Hybrid Mechanistic--Data-Driven Modeling Toward Neuro-Symbolic AI: What, Why, and How

    Jul 24, 2026Moein E. Samadi, Andreas SchuppertNeuro-Symbolic FrameworkArtificial Intelligence Models

  38. Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry

    Jul 22, 2026Jason Y. Hu, Ivan Higuera-Mendieta, Patrick Obin Sturm +1Atmospheric DynamicsWireless Foundation Models

  39. Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence

    Jul 20, 2026Katarzyna Filus, Sebastian PokucińskiMechanistic InterpretabilitySemantic Representations

  40. Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs

    Jul 18, 2026Sharath Naganna, Tanvir Ahmed Sijan, Uddipta KalitaLarge Language Models FailMechanistic Interpretability

  41. Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

    Jul 16, 2026Jihoon Hong, Julian Skifstad, Qiyue Dai +2World ModelsSteering

  42. Transcoders for Investigating Deception in Language Models

    Jul 16, 2026Darius Lim, Nathan Leow, Xin Wei ChiaDeceptionMechanistic Interpretability

  43. From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery

    Jul 14, 2026Ingmar Posner, Anson Lei, Bernhard SchölkopfModel DiscoveryMechanistic Interpretability

  44. Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

    Jul 13, 2026Zixiang Xu, Sixian Li, Huaxing Liu +4Llm-As-A-JudgeJudges

  45. Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

    Jul 9, 2026Bendegúz Váradi, Zoltán KmettyImproving Sparse AutoencodersBert-Based Models

  46. Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability

    Jul 9, 2026Amir AsiaeeInterpretabilityMechanistic Interpretability

  47. Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

    Jul 8, 2026Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang +1Large Language Model JailbreaksAdversarial Prompts

  48. Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

    Jul 8, 2026Pranav Sawant, Jakub KrejčíMechanistic InterpretabilityExplainable Artificial Intelligence

  49. Latent Programming Horizons in Coding Agents

    Jul 6, 2026André Silva, Han Tu, Martin MonperrusCoding AgentsMechanistic Interpretability

  50. Individual Parameters in Weight-Sparse Transformers Appear Interpretable

    Jul 3, 2026Arnau Marin-Llobet, Stefan HeimersheimMechanistic InterpretabilityModel Weights

  51. Towards Robustness against Typographic Attack with Training-free Concept Localization

    Jul 2, 2026Bohan Liu, Wenqian Ye, Guangzhi Xiong +3Contrastive Language-Image Pre-Training ModelVision Transformer

  52. Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits

    Jul 2, 2026Zhiren Gong, He Lu, Tiantong Wang +6AblationTransformer Architectures

  53. Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders

    Jul 1, 2026Zihao Qi, Christopher EarlsQubitMechanistic Interpretability