Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to AI agents. By fine-tuning on simulated prompt injections, SecAlign, a leading open defense, reports LLMs with good test-time robustness and negligible benign utility drop. By scaling up training and evaluations, however, we find that SecAlign actually suffers from significant utility degradation, especially in agentic tasks where the threat of prompt injection is prominent. Motivated by this, we propose Meta-SecAlign for utility-preserving defense by (1) randomized injection position during training to avoid shortcut learning and (2) self-generated responses as high-quality in-distribution training labels. Across general knowledge, instruction following, and agentic workflows (on tool-calling and web-navigation), Meta-SecAlign maintains almost all the undefended LLM's utility while achieving better overall security than SecAlign against various static and GCG adaptive attacks. Experiments use Llama-3.1-8B, Llama-3.3-70B, Llama-4-Scout, Qwen3-4B, and Qwen3.6-27B on 6 prompt injection benchmarks including AgentDojo, InjecAgent, WASP, and SEP. Below are links for the code (https://github.com/facebookresearch/Meta_SecAlign), Meta-SecAlign-70B (https://huggingface.co/facebook/Meta-SecAlign-70B), and Meta-SecAlign-8B (https://huggingface.co/facebook/Meta-SecAlign-8B) models.
Figures & tables
Label Annotator
-
SELF
SELF
davinci
GPT-5
Randomized Injection Position
-
No
Yes
Yes
Yes
MMLU ( ↑ )
86.3%
82.1%
85.9%
85.9%
85.8%
MMLU-Pro 5-shot ( ↑ )
67.7%
59.9%
67.6%
68.1%
67.3%
BBH 3-shot ( ↑ )
85.2%
80.0%
84.8%
85.3%
85.4%
GPQA Diamond ( ↑ )
50.0%
38.9%
48.0%
50.5%
49.5%
AlpacaEval2 Utility ( ↑ )
44.2%
43.2%
44.7%
40.6%
45.5%
Table 1: Randomized injection position improves utility while maintaining similar security, measured by attack success rate (ASR); see the third and fourth columns. Self-generated responses improve the utility-security trade-off (the fourth and fifth/sixth columns). Fine-tuning with low-quality labels ( text_davinci_003 ) or out-of-distribution labels ( GPT-5 ) performs worse than using labels from the initialization model, Llama-3.3-70B-Instruct , whose reference scores are in the grey column.
Llama-3.3-70B-Instruct
GPT
Gemini
Undef.
SecAlign
Ours
4o
5
2.5-FLH
3-Pro
MMLU ( ↑ )
86.3%
85.8%
85.9%
85.7%
93.5%
90.1%
93.9%
MMLU-Pro 5-shot ( ↑ )
67.7%
65.4%
67.6%
74.8%
87.1%
80.9%
90.0%
BBH 3-shot ( ↑ )
85.2%
84.5%
84.8%
83.0%
89.0%
73.5%
94.4%
GPQA Diamond ( ↑ )
50.0%
46.0%
48.0%
54.3%
85.1%
68.3%
91.0%
Table 2: Utility on General Knowledge Benchmarks.
Llama-3.3-70B-Instruct
GPT
Gemini
Undef.
SecAlign
Ours
4o
5
2.5-FLH
3-Pro
AlpacaEval2 Utility ( ↑ )
44.2%
38.7%
44.7%
56.4%
68.7%
44.6%
64.3%
SEP Utility ( ↑ )
62.1%
51.6%
60.4%
62.5%
76.8%
49.5%
70.1%
AlpacaFarm ASR ( ↓ )
95.7%
0.5%
0.5%
0%
1.0%
81.7%
0.5%
SEP ASR ( ↓ )
99.7%
12.4%
4.0%
37.4%
52.6%
76.9%
74.6%
TaskTracker ASR ( ↓ )
19.6%
0.2%
0.2%
0.6%
0.4%
1.1%
0.5%
Table 3: Utility and Attack Success Rate (ASR) on Instruction Following Benchmarks.
Llama-3.3-70B-Instruct
GPT
Gemini
Undef.
SecAlign
Ours
4o
5
2.5-FLH
3-Pro
AgentDojo Utility ( ↑ )
59.8%
6.2%
84.5%
79.4%
80.4%
63.9%
92.8%
- Utility w. Attack ( ↑ )
43.4%
7.9%
79.5%
67.4%
79.7%
52.6%
90.6%
WASP Utility ( ↑ )
62.2%
45.9%
59.5%
32.4%
59.5%
56.8%
59.5%
InjecAgent ASR ( ↓ )
53.8%
0%
0.5%
22.7%
0.5%
0.1%
0.2%
AgentDojo ASR ( ↓ )
14.7%
0%
1.9%
20.4%
0.2%
27.9%
2.3%
Table 4: Utility and Attack Success Rate (ASR) on Agentic Workflow Benchmarks.
Llama-3.1-8B-Instruct
Llama-3.3-70B-Instruct
Undef.
SecAlign
Ours
Undef.
SecAlign
Ours
MMLU ( ↑ )
72.0%
71.7%
71.7%
86.3%
85.8%
85.9%
MMLU-Pro 5-shot ( ↑ )
46.5%
45.9%
46.7%
67.7%
65.4%
67.6%
BBH 3-shot ( ↑ )
71.9%
71.2%
70.9%
85.2%
84.5%
84.8%
GPQA Diamond ( ↑ )
31.3%
30.8%
28.3%
50.0%
46.0%
48.0%
AlpacaEval2 Utility ( ↑ )
31.2%
30.7%
31.0%
44.2%
38.7%
44.7%
Table 5: Meta-SecAlign improves utility-security trade-off over SecAlign under adaptive GCG attack.
Qwen3-4B
Qwen3.6-27B
Llama-4-Scout
Undef.
Ours
Undef.
Ours
Undef.
Ours
MMLU ( ↑ )
70.7%
70.6%
90.0%
89.7%
85.9%
85.3%
MMLU-Pro ( ↑ )
64.6%
63.6%
84.8%
83.9%
71.7%
71.7%
BBH ( ↑ )
65.2%
67.2%
94.3%
92.8%
80.4%
77.6%
GPQA Diamond ( ↑ )
37.9%
37.9%
81.3%
79.3%
57.1%
54.0%
AlpacaEval2 Utility ( ↑ )
54.1%
55.1%
72.1%
70.7%
42.7%
43.0%
Table 6: Utility and Attack Success Rate (ASR) for more models before and after Meta-SecAlign.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Temperature
Output tokens
Total context
MMLU / MMLU-Pro
0
1,024
8,192
BBH
0
512
8,192
GPQA Diamond
0
2,048
8,192
AlpacaEval2 / AlpacaFarm
0
8,192
Model default
SEP / TaskTracker
0
8,192
Model default
InjecAgent
0
8,192
Model default
Appendix
Table 7: Decoding settings. Model default indicates that the default context limit is not overridden.
Llama-3.3-70B
GPT
Gemini
Undef.
Ours
4o
5
2.5-FLH
3-Pro
InjecAgent ASR ( ↓ )
86.0%
2.1%
36.9%
3.9%
3.5%
2.1%
AgentDojo Utility ( ↑ )
62.9%
79.4%
80.4%
83.5%
58.8%
93.8%
AgentDojo Utility w. Attack ( ↑ )
41.9%
77.1%
38.8%
81.3%
42.9%
89.0%
AgentDojo ASR ( ↓ )
23.0%
2.3%
43.2%
0.2%
30.7%
3.8%
Appendix
Table 8: InjecAgent and AgentDojo results without sandwich prompting.
Metric
Samples
Undefended
SecAlign
Meta-SecAlign
MMLU ( ↑ )
14,079
86.3 [85.72, 86.86]
85.8 [85.21, 86.37]
85.9 [85.32, 86.47]
MMLU-Pro ( ↑ )
12,032
67.7 [66.86, 68.54]
65.4 [64.54, 66.25]
67.6 [66.76, 68.44]
BBH ( ↑ )
6,511
85.2 [84.31, 86.05]
84.5 [83.60, 85.37]
84.8 [83.90, 85.66]
GPQA Diamond ( ↑ )
198
50.0 [42.83, 57.17]
46.0 [38.87, 53.17]
48.0 [40.85, 55.18]
AlpacaEval2 Utility ( ↑ )
805
44.2 [40.76, 47.73]
38.7 [35.38, 42.22]
44.7 [41.25, 48.23]
SEP Utility ( ↑ )
9,160
62.1 [61.09, 63.09]
51.6 [50.58, 52.63]
60.4 [59.39, 61.41]
Appendix
Table 9: Main results with approximate 95% confidence intervals for the Llama-3.3-70B-Instruct models. Each cell shows the reported percentage above its interval [lower,upper] .
Prompt injection remains a critical threat to LLM agents, yet existing defenses treat each task as a self-contained problem, independent of previous encounters. In practice, user requests are often underspecified: they describe the desired outcome without fully specifying acceptable behavior. An injection can exploit this ambiguity, causing the agent to complete the task in a way the user would reject. As the user's expectations become clearer through concrete cases, a defense should learn from each encounter and apply what it learns to the next. Inspired by adaptive immunity, we propose AgentAntibody, which equips LLM agents with a self-evolving immune system against prompt injection. AgentAntibody represents its evolving understanding of the user's security boundary as a persistent library of antibodies. At runtime, the library recognizes threats to this boundary and mounts corresponding immune responses. Across encounters, it evolves to strengthen the agent's immunity to future attacks. Extensive experiments across three benchmarks and four backbone LLMs show that, by learning the user's boundary through experience, AgentAntibody outperforms existing defenses in preventing harmful actions while preserving legitimate task completion, even when the harmful and legitimate actions are both compatible with the stated task.
Shihao Weng, Yang Feng, Xiaofei Xie +1
Nanjing University · Singapore Management University · Nanyang Technological University
Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain underexplored in realistic agentic settings. We present a comprehensive empirical evaluation of automated prompt injection attacks against LLM agents, adapting both white-box (GCG) and black-box (TAP) methods to the agentic setting within the AgentDojo framework. We evaluate across 80 task pairs spanning four domains and multiple models, and find that black-box optimization substantially outperforms gradient-based methods, a gap we attribute to GCG's optimization instability under reasonable compute budgets. We also find that TAP's effectiveness depends on the attacker model, as both general capability and safety tuning affect attack success--stronger models produce more effective injections, while safety-tuned attackers can refuse to generate adversarial prompts. Task-universal attacks transfer effectively to unseen tasks and out-of-distribution domains, but attacks optimized on smaller open-source models do not transfer to frontier models like GPT-5. These findings highlight automated prompt injection as a credible but model-dependent threat, with significant barriers remaining for model-agnostic exploitation.
LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on untrusted external content exposes them to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack agent behavior. Existing attacks rely on static payloads that cannot adapt to agent-specific defenses; even recent adaptive methods lack structured feedback to guide optimization. We introduce \oursys, a feedback-guided iterative framework that closes the loop between injection, diagnosis, and refinement: a rule-based diagnoser produces structured outcome labels with behavioral descriptions, and an LLM-based optimizer refines payloads conditioned on the full optimization history. A synthesis step generates new disguise seeds from failure patterns, enabling the strategy space to self-evolve. On AgentDojo and InjectAgent, \oursys substantially outperforms static baselines and existing adaptive methods across four victim models. Extension experiments on Claude Code, a production-grade coding agent with layered defenses, show that optimized payloads achieve full success on 5 of 9 targets; even those that resist full exploitation exhibit measurable improvement from iterative refinement. We further present a mechanistic analysis of IPI, identifying an attention-mediated threshold mechanism in mid-to-late layers; three causal interventions validate this finding and point to concrete defense directions.
Zixuan Chen, Jiaxiang Chen, Li Luo +4
1Shanghai Jiao Tong University · 2The University of Hong Kong