Causal Interventions in Language Models

Latest papers 317

All topics
CardsList
  1. Altered Thoughts, Altered Actions: Reasoning Chain as Control Surface for a Vision-Language-Action Policy

    Mar 13, 2026Tuan Duong Trinh, Basim Azam, Mohammed Ishaq Ansari +2Language-Conditioned Robot ManipulationVision-Language-Action Models

  2. From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions

    Mar 9, 2026Rishab Alagharu, Ishneet Sukhvinder Singh, Shaibi Shamsudeen +2Language Model SteeringLLM Refusal Behavior

  3. Patches of Nonlinearity: Instruction Vectors in Large Language Models

    Feb 8, 2026Irina Bigoulaeva, Jonas Rohweder, Subhabrata Dutta +1Transformer InterpretabilityLLM Interpretability

  4. Towards Isolated Interventions via Almost Orthogonal Features in Language Models

    Feb 4, 2026Moritz Miller, Florent Draye, Bernhard SchölkopfDisentangled Representation LearningRepresentation Learning

  5. Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models

    Feb 3, 2026Anish Sathyanarayanan, Aditya Nagarsekar, Aarush RathoreCoT FaithfulnessFaithfulness of Language Model Explanations

  6. There Is More to Refusal in Large Language Models than a Single Direction

    Feb 2, 2026Faaiz Joad, Majd Hawasly, Sabri Boughorbel +2LLM InterpretabilityLLM Refusal Behavior

  7. Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Reasoning

    Feb 2, 2026Yadong Wang, Haodong Chen, Yu Tian +3Sparse AutoencodersContinuous Latent Reasoning

  8. Cross-Lingual Activation Steering for Multilingual Language Models

    Jan 23, 2026Rhitabrat Pokharel, Ameeta Agrawal, Tanay NagarMultilingual Language ModelsLanguage Model Steering

  9. Tracing the Latent Threads: A Mechanistic Study of How LLMs Represent and Operationalize Race and Ethnicity Cues

    Jan 19, 2026Shiyue Hu, Ruizhe Li, Yanjun GaoLLM InterpretabilityMechanistic Interpretability

  10. To Copy or Not to Copy: Copying Is Easier to Induce Than Recall

    Jan 17, 2026Mehrdad Farahani, Franziska Penzkofer, Richard JohanssonKnowledge Conflicts in Language ModelsLLM Interpretability

  11. Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models

    Jan 12, 2026Zhenghao He, Guangzhi Xiong, Bohan Liu +2LLM InterpretabilityCoT Reasoning

  12. SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification

    Dec 17, 2025Hongbo Wang, AprilPyone MaungMaung, Isao EchizenMultimodal RobustnessLLM Safety

  13. LLMs Encode Harmfulness and Refusal Separately

    Jul 16, 2025Jiachen Zhao, Jing Huang, Zhengxuan Wu +2LLM InterpretabilityLLM Refusal Behavior

  14. When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning

    May 21, 2025Rongzhi Zhu, Yi Liu, Jiancheng Wang +6Overthinking in Language ModelsLLM Interpretability

  15. Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

    Mar 17, 2025Noah Y. Siegel, Nicolas Heess, Maria Perez-Ortiz +1Faithfulness of Language Model ExplanationsLLM Interpretability

  16. Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

    Date pendingElisabetta Rocchetti, Alfio FerraraLLM Refusal BehaviorCausal Interventions in Language Models