Activation Patching

Momentum

2 papers in the last four weeks, against 1 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 14

All topics
CardsList
  1. Social Circuits behind Multi-agent Echo Chambers

    Sep 28, 2026Chuiyang Meng, Wenlu Yu, Ming Tang +1Multi-Agent LLM SystemsMulti-Agent Systems

  2. StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions

    Sep 1, 2026Chao Gao, Haijiang Liu, Qiyuan Li +3Prompt SensitivityMultiple-Choice Question Answering

  3. The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

    Jun 25, 2026Sankaran Vaidyanathan, David Arbour, Aaron Mueller +2Causal Effect EstimationCausal Attribution

  4. Factual Retrieval in LLMs Is a Redundant, Distributed and Non-Contiguous Process

    Jun 19, 2026Hail Hochman, Natalie Shapira, Yoav GoldbergLLM InterpretabilityFactual Knowledge in Language Models

  5. Beyond Importance: Interchange-Sobol Sensitivity Reveals Task-Specific Content Channels in Transformer Components

    Jun 12, 2026Yifeng Guo, Jin-Hong Du, Xiang ChenMechanistic InterpretabilityCausal Interventions in Language Models

  6. When Attribution Patching Lies: Diagnosis and a Second-Order Correction

    Jun 5, 2026Luyang Zhang, Jialu WangGradient-Based AttributionLLM Interpretability

  7. Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability

    May 24, 2026David N. Olivieri, Antonio F. Pérez RodríguezTransformer InterpretabilityMechanistic Interpretability

  8. Patch-Effect Graph Kernels for LLM Interpretability

    May 7, 2026Ruben Fernandez-Boullon, David N. OlivieriGraph Representation LearningMechanistic Interpretability