Causal Interventions in Language Models

Latest papers 317

All topics
CardsList
  1. Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents

    Apr 30, 2026Kaituo Zhang, Zhen Xiong, Mingyu Zhong +4Tool-Augmented Language Model AgentsAgentic Reasoning

  2. DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

    Apr 30, 2026Lifan Zheng, Xue Yang, Jiawei Chen +6LLM InterpretabilityPersonality Modeling in Language Models

  3. Debiasing Reward Models via Causally Motivated Inference-Time Intervention

    Apr 30, 2026Kazutoshi Shinoda, Kosuke Nishida, Kyosuke NishidaReward ModelingLLM Alignment

  4. Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models

    Apr 29, 2026Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter +3LLM InterpretabilityInstruction Following

  5. What Suppresses Nash Equilibrium Play in Large Language Models? Mechanistic Evidence and Causal Control

    Apr 29, 2026Paraskevas V. Lekeas, Giorgos StamatopoulosNash EquilibriumGame Theory

  6. MoRFI: Monotonic Sparse Autoencoder Feature Identification

    Apr 29, 2026Dimitris Dimakopoulos, Shay B. Cohen, Ioannis KonstasLLM Hallucination MitigationLLM Interpretability

  7. Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors

    Apr 28, 2026Long Zhang, Zi-bo Qin, Wei-neng ChenKnowledge Conflicts in Language ModelsLLM Auditing

  8. Pruning via Causal Attribution Preserves Reasoning Performance in Large Language Models

    Apr 27, 2026Amogh Sheth, Biruk Assefa, Yi Wen Huang +2LLM PruningSelf-Attention

  9. When Chain-of-Thought Fails, the Solution Hides in the Hidden States

    Apr 25, 2026Houman Mehrafarin, Amit Parekh, Ioannis KonstasTransformer InterpretabilityCoT Reasoning

  10. Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization

    Apr 24, 2026Weixu Zhang, Ye Yuan, Changjiang Han +7LLM AlignmentLanguage Model Steering

  11. Dissociating Decodability and Causal Use in Bracket-Sequence Transformers

    Apr 24, 2026Aryan Sharma, Cutter Dawes, Shivam RavalTransformer InterpretabilitySelf-Attention

  12. Where Reasoning Breaks: Logic-Aware Path Selection by Controlling Logical Connectives in LLMs Reasoning Chains

    Apr 22, 2026Seunghyun Park, Yuanyuan LeiInference-Time SearchRL for Language Model Reasoning

  13. Separable Pathways for Causal Reasoning: How Architectural Scaffolding Enables Hypothesis-Space Restructuring in LLM Agents

    Apr 21, 2026John Alderete, Sebastian Benthal, Connie Xu +1Causal DiscoveryLLM Agents

  14. Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders

    Apr 21, 2026Het Patel, Tiejin Chen, Hua Wei +2LLM InterpretabilityLLM Reliability

  15. State Transfer Reveals Reuse in Controlled Routing

    Apr 20, 2026Yanzhen Lu, Zhicheng Qian, Muchen Jiang +1Prompt LearningTransfer Learning

  16. Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering

    Apr 19, 2026Li Zheng, Xin Zhang, Shuyi He +5Language Model SteeringLLM Interpretability

  17. One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them

    Apr 18, 2026Ali Holmov, Paul Youssef, Nandi Schoots +1Knowledge EditingTransformer Interpretability

  18. CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification

    Apr 16, 2026Yian Wang, Yuen Chen, Agam Goyal +1Causal Reasoning in Language ModelsLLM Safety

  19. Mechanistic Decoding of Cognitive Constructs in Large Language Models

    Apr 16, 2026Yitong Shou, Manhao GuanLLM InterpretabilityCausal Interventions in Language Models