Causal Interventions in Language Models

Latest papers 317

All topics
CardsList
  1. BeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language Models

    Oct 8, 2026Shuai Guo, Yidong CuiBelief Updating in LLMsLLM Auditing

  2. Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning

    Oct 7, 2026Wei Zhai, Xiang Liu, Qiang Huang +6LLM UnlearningPrivacy Leakage in Language Models

  3. The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models

    Oct 7, 2026Zhe Yu, Wenpeng Xing, Yunzhao Wei +4LLM InterpretabilityRetrieval-Augmented Generation

  4. TrustMI: Causally controlling how assistants trust their users

    Oct 5, 2026Théo Lasnier, Romain Froger, Maxence Lasbordes +1LLM SafetyTrust in AI

  5. Readable Before Actionable: Causal Tracing of Indirect Prompt Injection

    Oct 4, 2026Zhe Yu, Wenpeng Xing, Xingxing Yang +1Causal Reasoning in Language ModelsIndirect Prompt Injection

  6. Mind the Gaps: From Failure Attribution to Closed-Form Repair of Code Language Models

    Oct 4, 2026Jian Gu, Hongyu Zhang, Chunyang Chen +1Causal Interventions in Language ModelsCode Language Models

  7. Usage-Modulated Sentiment Representations in Large Language Models

    Oct 4, 2026Hongfei Du, Jiacheng Shi, Yanfu Zhang +2Causal Interventions in Language ModelsRepresentation Geometry in Language Models

  8. Clinical Concept Centers in LLMs

    Oct 2, 2026Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal +2Mechanistic InterpretabilityCausal Interventions in Language Models

  9. Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining

    Oct 1, 2026Shengye Tao, Yinzhu Cheng, Haihua XieLLM InterpretabilityCausal Interventions in Language Models

  10. When Do Biological Reasoning Models Use Their Biological Inputs?

    Oct 1, 2026Ada Fang, Nikitha Thoduguli, Lukas Fesser +3Faithfulness of Language Model ExplanationsCausal Interventions in Language Models

  11. Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints

    Sep 30, 2026Peter Nutter, Dani Roytburg, Clément Dumas +2Fine-TuningLLM Alignment

  12. Semifactual Credit-Augmented Policy Optimization

    Sep 30, 2026Junshu Pan, Zhizhang Fu, Shulin Huang +5Prompt SensitivityRL for Language Model Reasoning

  13. Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering

    Sep 30, 2026Jiale Dai, Hongcan Deng, Liuxian Ma +2Disentangled Representation LearningLLM Alignment

  14. When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents

    Sep 30, 2026Shuyao Xiao, Shengling Wang, Xuan Chen +7Long-Horizon Agent EvaluationCausal Interventions in Language Models

  15. Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis

    Sep 30, 2026Xia Hu, Brian Potetz, Chun-Ta Lu +6LLM EvaluationCounterfactual Evaluation of VLMs

  16. Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

    Sep 29, 2026Zehao Jin, Ruixuan Deng, Junran WangTransformerLLM Fine-Tuning

  17. Causal and Interpretable Structures in LLM Compositional Tasks

    Sep 28, 2026Gurbir Arora, Toni J. B. Liu, Jiajun Bao +2LLM InterpretabilityCausal Interventions in Language Models

  18. Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

    Sep 28, 2026Xu Wang, Difan Zou, Xuansheng WuLLM SycophancyLanguage Model Safety Evaluation

  19. Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices

    Sep 28, 2026Lucio La Cava, Andrea TagarelliMoral Reasoning in Language ModelsLLM Auditing

  20. Causal Routing for Unlearning

    Sep 28, 2026Bardh Prenkaj, Andrea D'Angelo, Davide Mottin +4Causal Interventions in Language ModelsMachine Unlearning

  21. Steering Language Model Goals with Value Transplant

    Sep 28, 2026Pengcheng Jiang, Fabien RogerLanguage Model SteeringLanguage Model-Based Control