Causal Interventions in Language Models

Latest papers 317

All topics
CardsList
  1. Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination

    May 11, 2026Yangneng Chen, Junlin Li, Weijun Yao +4Hallucination in Language ModelsVLM Hallucination

  2. Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models

    May 10, 2026Cameron Berg, Roshni LullaLLM InterpretabilityPersonality Modeling in Language Models

  3. Causal state binding predicts action control in language agents

    May 10, 2026Xiao JiaLanguage Model-Based ControlLLM Agent Evaluation

  4. Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal

    May 10, 2026Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang +2LLM InterpretabilityCausal Interventions in Language Models

  5. Decomposing and Steering Functional Metacognition in Large Language Models

    May 9, 2026Yanshi Li, Xueru Bai, Shuman Liu +2LLM EvaluationLLM Interpretability

  6. A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models

    May 8, 2026Hamid Kazemi, Atoosa Chegini, Maria SafiLLM InterpretabilityAdversarial Attacks on LLMs

  7. How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

    May 8, 2026Michael Li, Nishant SubramaniCircuit DiscoveryMechanistic Interpretability

  8. Tool Calling is Linearly Readable and Steerable in Language Models

    May 8, 2026Zekun Wu, Ze Wang, Seonglae Cho +4AI Agent ReliabilityTool-Augmented Language Model Agents

  9. Inference Time Causal Probing in LLMs

    May 8, 2026Sadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash +1Causal Interventions in Language ModelsActivation Steering

  10. Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts

    May 8, 2026Yi-Chang Chen, Feng-Ting Liao, Da-shan Shiu +1CoT ReasoningCausal Interventions in Language Models

  11. Crafting Reversible SFT Behaviors in Large Language Models

    May 7, 2026Yuping Lin, Pengfei He, Yue Xing +5Supervised Fine-TuningLanguage Model Steering

  12. Patch-Effect Graph Kernels for LLM Interpretability

    May 7, 2026Ruben Fernandez-Boullon, David N. OlivieriGraph Representation LearningMechanistic Interpretability

  13. Don't Lose Focus: Activation Steering via Key-Orthogonal Projections

    May 7, 2026Haoyan Luo, Mateo Espinosa Zarlenga, Mateja JamnikLanguage Model SteeringSelf-Attention

  14. Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions

    May 7, 2026Yuntai Bao, Qinfeng Li, Xinyan Yu +6Language Model SteeringCausal Interventions in Language Models

  15. Causal Probing for Internal Visual Representations in Multimodal Large Language Models

    May 7, 2026Zehao Deng, Tianjie Ju, Zheng Wu +5Multimodal Large Language ModelsVLM Interpretability

  16. Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior

    May 6, 2026Daniel Wurgaft, Can Rager, Matthew Kowal +13Neural Representation GeometryManifold Learning

  17. Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models

    May 6, 2026Quintin Pope, Ajay Hayagreeve Balaji, Jacques Thibodeau +1Language Model Generation EvaluationLLM Auditing

  18. Steer Like the LLM: Activation Steering that Mimics Prompting

    May 5, 2026Geert Heyman, Frederik VandeputteLanguage Model SteeringLLM Prompting

  19. Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models

    May 5, 2026Chenchen Yuan, Zheyu Zhang, Gjergji KasneciLLM AlignmentMoral Reasoning in Language Models

  20. Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes

    May 4, 2026Michael A. Riegler, Birk Sebastian Frostelid Torpmann-HagenSparse AutoencodersMechanistic Interpretability

  21. Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation

    May 4, 2026Francesco Sovrano, Gabriele Dominici, Marc LangheinrichLLM InterpretabilityMechanistic Interpretability

  22. Compared to What? Baselines and Metrics for Counterfactual Prompting

    May 1, 2026Zihao Yang, Mosh Levy, Yoav Goldberg +1Counterfactual EvaluationCausal Interventions in Language Models

  23. Attention Is Where You Attack

    Apr 30, 2026Aviral Srivastava, Sourav PandaSelf-AttentionAdversarial Attacks