Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to AI agents. By fine-tuning on simulated prompt injections, SecAlign, a leading open defense, reports LLMs with good test-time robustness and negligible benign utility drop. By scaling up training and evaluations, however, we find that SecAlign actually suffers from significant utility degradation, especially in agentic tasks where the threat of prompt injection is prominent. Motivated by this, we propose Meta-SecAlign for utility-preserving defense by (1) randomized injection position during training to avoid shortcut learning and (2) self-generated responses as high-quality in-distribution training labels. Across general knowledge, instruction following, and agentic workflows (on tool-calling and web-navigation), Meta-SecAlign maintains almost all the undefended LLM's utility while achieving better overall security than SecAlign against various static and GCG adaptive attacks. Experiments use Llama-3.1-8B, Llama-3.3-70B, Llama-4-Scout, Qwen3-4B, and Qwen3.6-27B on 6 prompt injection benchmarks including AgentDojo, InjecAgent, WASP, and SEP. Below are links for the code (https://github.com/facebookresearch/Meta_SecAlign), Meta-SecAlign-70B (https://huggingface.co/facebook/Meta-SecAlign-70B), and Meta-SecAlign-8B (https://huggingface.co/facebook/Meta-SecAlign-8B) models.
Figures & tables
Label Annotator
-
SELF
SELF
davinci
GPT-5
Randomized Injection Position
-
No
Yes
Yes
Yes
MMLU ( ↑ )
86.3%
82.1%
85.9%
85.9%
85.8%
MMLU-Pro 5-shot ( ↑ )
67.7%
59.9%
67.6%
68.1%
67.3%
BBH 3-shot ( ↑ )
85.2%
80.0%
84.8%
85.3%
85.4%
GPQA Diamond ( ↑ )
50.0%
38.9%
48.0%
50.5%
49.5%
AlpacaEval2 Utility ( ↑ )
44.2%
43.2%
44.7%
40.6%
45.5%
Table 1: Randomized injection position improves utility while maintaining similar security, measured by attack success rate (ASR); see the third and fourth columns. Self-generated responses improve the utility-security trade-off (the fourth and fifth/sixth columns). Fine-tuning with low-quality labels ( text_davinci_003 ) or out-of-distribution labels ( GPT-5 ) performs worse than using labels from the initialization model, Llama-3.3-70B-Instruct , whose reference scores are in the grey column.
Llama-3.3-70B-Instruct
GPT
Gemini
Undef.
SecAlign
Ours
4o
5
2.5-FLH
3-Pro
MMLU ( ↑ )
86.3%
85.8%
85.9%
85.7%
93.5%
90.1%
93.9%
MMLU-Pro 5-shot ( ↑ )
67.7%
65.4%
67.6%
74.8%
87.1%
80.9%
90.0%
BBH 3-shot ( ↑ )
85.2%
84.5%
84.8%
83.0%
89.0%
73.5%
94.4%
GPQA Diamond ( ↑ )
50.0%
46.0%
48.0%
54.3%
85.1%
68.3%
91.0%
Table 2: Utility on General Knowledge Benchmarks.
Llama-3.3-70B-Instruct
GPT
Gemini
Undef.
SecAlign
Ours
4o
5
2.5-FLH
3-Pro
AlpacaEval2 Utility ( ↑ )
44.2%
38.7%
44.7%
56.4%
68.7%
44.6%
64.3%
SEP Utility ( ↑ )
62.1%
51.6%
60.4%
62.5%
76.8%
49.5%
70.1%
AlpacaFarm ASR ( ↓ )
95.7%
0.5%
0.5%
0%
1.0%
81.7%
0.5%
SEP ASR ( ↓ )
99.7%
12.4%
4.0%
37.4%
52.6%
76.9%
74.6%
TaskTracker ASR ( ↓ )
19.6%
0.2%
0.2%
0.6%
0.4%
1.1%
0.5%
Table 3: Utility and Attack Success Rate (ASR) on Instruction Following Benchmarks.
Llama-3.3-70B-Instruct
GPT
Gemini
Undef.
SecAlign
Ours
4o
5
2.5-FLH
3-Pro
AgentDojo Utility ( ↑ )
59.8%
6.2%
84.5%
79.4%
80.4%
63.9%
92.8%
- Utility w. Attack ( ↑ )
43.4%
7.9%
79.5%
67.4%
79.7%
52.6%
90.6%
WASP Utility ( ↑ )
62.2%
45.9%
59.5%
32.4%
59.5%
56.8%
59.5%
InjecAgent ASR ( ↓ )
53.8%
0%
0.5%
22.7%
0.5%
0.1%
0.2%
AgentDojo ASR ( ↓ )
14.7%
0%
1.9%
20.4%
0.2%
27.9%
2.3%
Table 4: Utility and Attack Success Rate (ASR) on Agentic Workflow Benchmarks.
Llama-3.1-8B-Instruct
Llama-3.3-70B-Instruct
Undef.
SecAlign
Ours
Undef.
SecAlign
Ours
MMLU ( ↑ )
72.0%
71.7%
71.7%
86.3%
85.8%
85.9%
MMLU-Pro 5-shot ( ↑ )
46.5%
45.9%
46.7%
67.7%
65.4%
67.6%
BBH 3-shot ( ↑ )
71.9%
71.2%
70.9%
85.2%
84.5%
84.8%
GPQA Diamond ( ↑ )
31.3%
30.8%
28.3%
50.0%
46.0%
48.0%
AlpacaEval2 Utility ( ↑ )
31.2%
30.7%
31.0%
44.2%
38.7%
44.7%
Table 5: Meta-SecAlign improves utility-security trade-off over SecAlign under adaptive GCG attack.
Qwen3-4B
Qwen3.6-27B
Llama-4-Scout
Undef.
Ours
Undef.
Ours
Undef.
Ours
MMLU ( ↑ )
70.7%
70.6%
90.0%
89.7%
85.9%
85.3%
MMLU-Pro ( ↑ )
64.6%
63.6%
84.8%
83.9%
71.7%
71.7%
BBH ( ↑ )
65.2%
67.2%
94.3%
92.8%
80.4%
77.6%
GPQA Diamond ( ↑ )
37.9%
37.9%
81.3%
79.3%
57.1%
54.0%
AlpacaEval2 Utility ( ↑ )
54.1%
55.1%
72.1%
70.7%
42.7%
43.0%
Table 6: Utility and Attack Success Rate (ASR) for more models before and after Meta-SecAlign.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Temperature
Output tokens
Total context
MMLU / MMLU-Pro
0
1,024
8,192
BBH
0
512
8,192
GPQA Diamond
0
2,048
8,192
AlpacaEval2 / AlpacaFarm
0
8,192
Model default
SEP / TaskTracker
0
8,192
Model default
InjecAgent
0
8,192
Model default
Appendix
Table 7: Decoding settings. Model default indicates that the default context limit is not overridden.
Llama-3.3-70B
GPT
Gemini
Undef.
Ours
4o
5
2.5-FLH
3-Pro
InjecAgent ASR ( ↓ )
86.0%
2.1%
36.9%
3.9%
3.5%
2.1%
AgentDojo Utility ( ↑ )
62.9%
79.4%
80.4%
83.5%
58.8%
93.8%
AgentDojo Utility w. Attack ( ↑ )
41.9%
77.1%
38.8%
81.3%
42.9%
89.0%
AgentDojo ASR ( ↓ )
23.0%
2.3%
43.2%
0.2%
30.7%
3.8%
Appendix
Table 8: InjecAgent and AgentDojo results without sandwich prompting.
Metric
Samples
Undefended
SecAlign
Meta-SecAlign
MMLU ( ↑ )
14,079
86.3 [85.72, 86.86]
85.8 [85.21, 86.37]
85.9 [85.32, 86.47]
MMLU-Pro ( ↑ )
12,032
67.7 [66.86, 68.54]
65.4 [64.54, 66.25]
67.6 [66.76, 68.44]
BBH ( ↑ )
6,511
85.2 [84.31, 86.05]
84.5 [83.60, 85.37]
84.8 [83.90, 85.66]
GPQA Diamond ( ↑ )
198
50.0 [42.83, 57.17]
46.0 [38.87, 53.17]
48.0 [40.85, 55.18]
AlpacaEval2 Utility ( ↑ )
805
44.2 [40.76, 47.73]
38.7 [35.38, 42.22]
44.7 [41.25, 48.23]
SEP Utility ( ↑ )
9,160
62.1 [61.09, 63.09]
51.6 [50.58, 52.63]
60.4 [59.39, 61.41]
Appendix
Table 9: Main results with approximate 95% confidence intervals for the Llama-3.3-70B-Instruct models. Each cell shows the reported percentage above its interval [lower,upper] .