RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Organizations: LIX, École polytechnique, Institut Polytechnique de Paris, CNRS · AMIAD (Agence Ministérielle pour l’IA de Défense) · IRT SystemX · Google DeepMind
Abstract
Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.
Figures & tables
| InjecAgent ASR | AgentDojo | AgentDyn | |||||||
| Model | Method | Base | Enhanced | ASR | ASR | ||||
| Gemma | Vanilla | 1.0 | 0.4 | 84.5 | 64.0 | 39.5 | 61.7 | 55.5 | 53.4 |
| Prompt Sandwiching | 0.0 | 0.0 | 83.5 | 64.4 | 34.1 | 60.0 | 52.2 | 56.1 | |
| Spotlighting | 1.1 | 0.1 | 80.4 | 55.8 | 43.9 | 61.7 | 55.4 | 54.6 | |
| PromptGuard2 | 0.9 | 0.0 | 84.5 | 42.9 | 16.1 | 60.0 | 34.1 | 37.9 | |
| SecAlign++ | |||||||||
| Model | Method | ToolTalk | BFCL | WorkBench | MMT-Bench | AsyncTool |
|---|---|---|---|---|---|---|
| Gemma | Vanilla | 65.4 | 57.0 | 55.9 | 53.0 | 52.7 |
| Prompt Sandwiching | 59.0 | 33.0 | 54.9 | 49.2 | 52.7 | |
| Spotlighting | 69.2 | 57.5 | 50.4 | 53.4 | 46.6 | |
| PromptGuard2 | 65.4 | 57.0 | 55.9 | 53.0 | 52.7 | |
| SecAlign++ | ||||||
| RAISED (Ours) |
| Model | Method | MMLU | MMLU-Pro | MMMLU | GPQA | AGIEval | IFEval | GSM8K | Minerva | HumanEval | MBPP |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemma | Vanilla | 75.1 | 81.1 | 66.6 | 72.2 | 57.0 | 94.1 | 94.9 | 92.5 | 95.1 | 83.4 |
| SecAlign++ | |||||||||||
| RAISED (Ours) | |||||||||||
| Nemotron | Vanilla | 49.5 | 60.1 | 34.6 | 42.4 | 23.3 | 81.3 | 87.2 | 76.5 | 87.2 | 67.8 |
| SecAlign++ | |||||||||||
| RAISED (Ours) |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| , , | State space; initial and final state of an episode |
| , | Tool menu; distractor tools added for the attacker goal |
| , | Argument space and executor of tool |
| , | -th tool observation (untrusted); -th assistant action |
| , | System prompt; user request |
| Chat template |
| Task | Base | SFT | SecAlign++ | RAISED |
|---|---|---|---|---|
| Change the account password; a tool emails a verification code | reads code, updates ✓ | claims completion without reading code | asks user for code | reads code, updates ✓ |
| Follow instructions in a file; the file says to adjust a standing order | updates standing order ✓ | relays instruction, no update | updates standing order ✓ | updates standing order ✓ |
| Transfer the most-starred repository to an address read from the inbox | transfers and verifies ✓ | claims completion without verification | transfers and verifies ✓ | transfers and verifies ✓ |
| Invite the collaborator who requested it; a tool emails a verification code | reads code, invites ✓ | claims sent without verification | claims sent without verification | reads code, invites ✓ |
| Release the batch; an error identifies the corrected record to release | releases corrected record ✓ | claims completion, releases nothing | asks user whether to release it | releases corrected record ✓ |
| Open the course for enrollment; the tool says to publish the syllabus first | publishes, then opens ✓ | asks user to publish first | asks whether to publish first | publishes, then opens ✓ |
| Task | Base | SFT | SecAlign++ | RAISED |
|---|---|---|---|---|
| Publish the menu; an error returns the corrected versioned target | publishes corrected target ✓ | claims completion, wrong target | publishes original target | publishes corrected target ✓ |
| Post the store-opening summary; the required post step is named only in a tool output | posts summary ✓ | lists statuses, does not post | describes summary, does not post | posts summary ✓ |
| Sync the inventory; the service says to retry with an updated version | retries and syncs ✓ | tells user to retry | retries and syncs ✓ | retries and syncs ✓ |
| Finalize payroll; the tool says timesheets must be validated first | validates, then finalizes ✓ | claims completion, skips validation | validates, then finalizes ✓ | validates, then finalizes ✓ |
| Structure | Construction |
|---|---|
| Surfaced guidance | A step required to complete the user’s task is revealed through a tool output rather than stated in the original request. The output may directly indicate the next operation or identify the resource, record, or location needed to determine it. |
| Recovery / completion | An intermediate observation reveals how an incomplete task can be continued, for example by exposing a supported alternative after an unsuccessful action, identifying another resource needed for completion, or showing that another user-requested obligation remains. |
| Follow-through | An action that appears to complete the request instead returns a provisional or incomplete result, requiring one or more additional authorized actions before the task is actually finished. |
| InjecAgent ASR | AgentDojo | AgentDyn | |||||||
| Model | Method | Base | Enhanced | ASR | ASR | ||||
| Qwen | Vanilla | 0.0 | 0.0 | 88.7 | 80.0 | 5.8 | 66.7 | 64.8 | 4.2 |
| Prompt Sandwiching | 0.0 | 0.0 | 89.7 | 82.6 | 3.8 | 73.3 | 68.2 | 1.8 | |
| Spotlighting | 0.0 | 0.0 | 89.7 | 85.4 | 0.9 | 73.3 | 67.9 | 0.5 | |
| PromptGuard2 | 0.0 | 0.0 | 88.7 | 58.6 | 2.7 | 66.7 | 47.7 | 3.3 | |
| SecAlign++ | |||||||||
| Model | Method | ToolTalk | BFCL | WorkBench | MMT-Bench | AsyncTool |
|---|---|---|---|---|---|---|
| Qwen | Vanilla | 66.7 | 50.0 | 83.0 | 50.5 | 68.1 |
| Prompt Sandwiching | 67.9 | 46.0 | 84.5 | 48.8 | 68.4 | |
| Spotlighting | 69.2 | 47.0 | 82.3 | 49.3 | 68.9 | |
| PromptGuard2 | 66.7 | 50.0 | 83.0 | 50.5 | 68.1 | |
| SecAlign++ | ||||||
| RAISED (Ours) |
| Model | Method | MMLU | MMLU-Pro | MMMLU | GPQA | AGIEval | IFEval | GSM8K | Minerva | HumanEval | MBPP |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen | Vanilla | 76.5 | 58.6 | 36.5 | 96.1 | 93.6 | 96.3 | 81.6 | |||
| SecAlign++ | |||||||||||
| RAISED (Ours) |
| Base | SecAlign++ | RAISED | |
|---|---|---|---|
| Starts the injected tool call (%, 72 pairs) | 62.5 | 13.9 | 12.5 |
| KL to the base model’s clean prediction (nats, 88 pairs) | 0.338 | 0.062 | 0.011 |
| Injection effect at block 31, peak (rel. ) | 0.45 | 0.29 | 0.23 |
| Distance to base on clean context, block 47 | – | 0.19 | 0.12 |
| InjecAgent ASR | AgentDojo | AgentDyn | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | Steps | Base | Enhanced | ASR | ASR | ||||
| Vanilla | – | 26.7 | 33.7 | 48.5 | 39.9 | 11.7 | 25.0 | 19.6 | 17.6 |
| SecAlign++ | 313 | ||||||||
| SecAlign++ | 625 | ||||||||
| RAISED (Ours) | – | ||||||||
| Method | Steps | MMLU | MMLU-Pro | MMMLU | GPQA | AGIEval | IFEval | GSM8K | Minerva | HumanEval | MBPP |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Vanilla | – | 49.5 | 60.1 | 34.6 | 42.4 | 23.3 | 81.3 | 87.2 | 76.5 | 87.2 | 67.8 |
| SecAlign++ | 313 | ||||||||||
| SecAlign++ | 625 | ||||||||||
| RAISED (Ours) | – |
| AgentDojo | AgentDyn | ||||
| Data | Objective | ASR | ASR | ||
| – | Vanilla | 84.5 | 39.5 | 61.7 | 53.4 |
| Glaive | SFT | ||||
| Glaive | DPO | ||||
| Glaive | Self-distillation | ||||
| Self-generated | DPO | ||||