cs.CROct 5, 2026

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

Authors: Mohamed Dhouib, Clement Elliker, Alexi Canesse, Maël Jenny, Lucas-Andrei Thil, Mahammed El Sharkawy, Sonia Vanier, Elie Bursztein

Organizations: LIX, École polytechnique, Institut Polytechnique de Paris, CNRS · AMIAD (Agence Ministérielle pour l’IA de Défense) · IRT SystemX · Google DeepMind

Abstract

Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents

    Jul 3, 2025Sizhe Chen, Arman Zharmagambetov, David Wagner +1Indirect Prompt InjectionGradient-Based Attacks

  2. Assessing Automated Prompt Injection Attacks in Agentic Environments

    Jun 9, 2026David Hofer, Edoardo Debenedetti, Florian TramèrIndirect Prompt InjectionLarge Language Model Agents

  3. AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection

    Aug 4, 2026Shihao Weng, Yang Feng, Xiaofei Xie +1Multi-Llm AgentsPrompt Engineering