Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI splits an attack objective into incomplete fragments, embeds them in external content retrieved through the agent's tools, and uses a reconstruction cue to prompt the agent to combine them. It then iteratively refines the fragments and cue with OpenEvolve using graded scoring and natural-language execution feedback from the target agent. Empirically, AdaLCPI achieves higher attack success than strong adaptive baselines, reaching 61.4% macro-average ASR compared with 32.8% for Trojan Hippo-style and 30.0% for AgentVigil. Safety evaluations should therefore test whether agents remain robust when harmful objectives must be reconstructed from incomplete fragments.
Figures & tables
Figure 1: AdaLCPI makes long-context fragmentation adaptive. An attack objective is split into two incomplete fragments and a reconstruction cue embedded in long external content. An adaptive loop then uses graded scoring and natural-language execution feedback to refine them.
Baselines (ASR)
AdaLCPI (ASR)
Δ
Utility
Model
TAP
Hippo
AgentVigil
IterInject
LCF
2k
4k
8k
16k
AdaLCPI-8k
GPT-4.1
66.7
56.8
54.3
64.2
2.5
51.9
53.1
58.0
66.7
0.0
42.0
GPT-5.1
17.3
46.9
33.3
34.6
3.7
55.6
59.3
65.4
49.4
+18.5
53.1
GPT-5.6-Luna
2.5
1.2
0.0
1.2
0.0
8.6
8.6
8.6
14.8
+12.3
66.7
Qwen-3.6-27B
1.2
9.9
8.6
17.3
4.9
65.4
71.6
71.6
72.8
+62.9
53.1
Ministral-3-14B
84.0
54.3
51.9
56.8
3.7
39.5
39.5
61.7
63.0
-21.0
25.9
Table 1: AdaLCPI outperforms both non-adaptive long-context fragmentation and adaptive explicit-instruction attacks. We report attack success rate (ASR, %) and benign task completion under attack over 81 task–objective pairs per model. The Trojan Hippo-style baseline uses the same adaptive search, graded scoring, and natural-language execution feedback as AdaLCPI but keeps the malicious instruction explicit and uses no added filler. LCF denotes long-context fragmentation without adaptive optimization. TAP represents search based approaches, while AgentVigil and IterInject represent agentic attacks. Δ compares the best AdaLCPI result with the stronger of TAP, Trojan Hippo-style, AgentVigil and IterInject. Utility reports benign task completion under AdaLCPI-8k.
TAP
Hippo
AgentVigil
IterInject
LCF
Ada-8k
Δ
Environment
Email
24.3
31.7
37.0
30.2
1.6
69.8
+38.1
GitHub
37.0
50.8
40.2
49.2
13.8
59.3
+8.5
Slack
17.5
15.9
12.7
18.5
0.0
55.0
+39.2
Attack category
Data Exfiltration
26.5
38.6
43.4
39.2
3.2
66.1
+27.5
Table 2: Attack success rate (ASR, %) averaged across seven models by environment and attack category. The Trojan Hippo-style baseline uses no added filler. Δ is AdaLCPI-8k minus the Trojan Hippo-style baseline, in percentage points. Averages and differences are computed before rounding.
Email
GitHub
Slack
Model
c
f1+c
f2+c
c
f1+c
f2+c
c
f1+c
f2+c
GPT-5.1
0.0
0.0
15.0
0.0
0.0
3.5
0.0
0.0
0.0
Qwen-3.6-27B
0.0
6.3
4.8
0.0
3.2
7.9
0.0
0.0
2.1
Gemma-4-31B
0.0
0.0
29.3
0.0
13.0
27.5
0.0
0.0
23.1
Muse-Glimmer-30B
0.0
5.3
24.0
0.0
1.7
10.0
0.0
0.0
2.4
Table 3: The reconstruction cue alone does not reproduce successful AdaLCPI-8k attacks. We report attack success rate (ASR, %) after removing fragments from attacks that originally succeeded. Here, c retains only the reconstruction cue, while f1+c and f2+c retain one fragment and the cue.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Max iterations
20
Candidates generated per iteration
6
Parent sample size per iteration
4
Elite ratio (elite/non-elite parent split)
0.5
Early-stop patience (iterations w/o improvement)
8
Initial seed candidates
1
Appendix
Table 4: Search hyperparameters used for AdaLCPI. The quality-diversity archive uses payload length and normalized edit distance from the initial seed x0 as its two feature dimensions.
Role
Model
Temperature
Max generation
Target agent
–
0.6
–
Generator
GPT-4.1
1.0
–
Mutator
GPT-4.1
0.8
8192
Binary Judge (stage 1)
GPT-4o
0.0
–
Graded Judge (stage 2)
GPT-4o
1.0
–
Appendix
Table 5: Models and decoding settings used by the AdaLCPI components. The target agent is the target model under evaluation and therefore varies across experiments.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
55.6
33.3
77.8
55.6
100.0
55.6
100.0
85.2
22.2
11.1
55.6
29.6
GPT-5.1
100.0
44.4
88.9
77.8
66.7
33.3
66.7
55.6
0.0
0.0
22.2
7.4
GPT-5.6-Luna
0.0
0.0
0.0
0.0
11.1
0.0
0.0
3.7
0.0
0.0
0.0
0.0
Qwen-3.6-27B
0.0
0.0
0.0
0.0
44.4
0.0
33.3
25.9
0.0
11.1
0.0
3.7
Ministral-3-14B
22.2
0.0
44.4
22.2
77.8
66.7
88.9
77.8
55.6
77.8
55.6
63.0
Appendix
Table 6: Trojan Hippo-style attack success rate (ASR, %) by environment and attack-objectives. This baseline uses adaptive search while keeping the malicious instruction explicit.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
100.0
77.8
100.0
92.6
100.0
66.7
100.0
88.9
0.0
0.0
55.6
18.5
GPT-5.1
0.0
0.0
22.2
7.4
44.4
0.0
77.8
40.7
0.0
0.0
11.1
3.7
GPT-5.6-Luna
0.0
0.0
0.0
0.0
0.0
0.0
22.2
7.4
0.0
0.0
0.0
0.0
Qwen-3.6-27B
0.0
0.0
0.0
0.0
0.0
0.0
11.1
3.7
0.0
0.0
0.0
0.0
Ministral-3-14B
66.7
66.7
66.7
66.7
100.0
66.7
100.0
88.9
100.0
100.0
88.9
96.3
Appendix
Table 7: TAP attack success rate (ASR, %) by environment and attack-objectives.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
100.0
55.6
100.0
85.2
100.0
0.0
100.0
66.7
22.2
0.0
11.1
11.1
GPT-5.1
66.7
33.3
66.7
55.6
44.4
0.0
88.9
44.4
0.0
0.0
0.0
0.0
GPT-5.6-Luna
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Qwen-3.6-27B
44.4
11.1
11.1
22.2
0.0
0.0
0.0
0.0
0.0
11.1
0.0
3.7
Ministral-3-14B
55.6
11.1
22.2
29.6
100.0
44.4
88.9
77.8
77.8
33.3
33.3
48.1
Appendix
Table 8: AgentVigil attack success rate (ASR, %) by environment and attack-objectives.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
100.0
33.3
77.8
70.4
100.0
44.4
100.0
81.5
22.2
0.0
100.0
40.7
GPT-5.1
33.3
22.2
66.7
40.7
66.7
33.3
66.7
55.6
0.0
0.0
22.2
7.4
GPT-5.6-Luna
0.0
0.0
0.0
0.0
0.0
11.1
0.0
3.7
0.0
0.0
0.0
0.0
Qwen-3.6-27B
44.4
22.2
22.2
29.6
22.2
0.0
22.2
14.8
11.1
0.0
11.1
7.4
Ministral-3-14B
33.3
0.0
44.4
25.9
100.0
55.6
88.9
81.5
100.0
66.7
22.2
63.0
Appendix
Table 9: IterInject attack success rate (ASR, %) by environment and attack-objectives.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
11.1
0.0
0.0
3.7
0.0
11.1
0.0
3.7
0.0
0.0
0.0
0.0
GPT-5.1
0.0
0.0
0.0
0.0
0.0
11.1
22.2
11.1
0.0
0.0
0.0
0.0
GPT-5.6-Luna
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Qwen-3.6-27B
0.0
0.0
0.0
0.0
11.1
22.2
11.1
14.8
0.0
0.0
0.0
0.0
Ministral-3-14B
0.0
0.0
0.0
0.0
11.1
22.2
0.0
11.1
0.0
0.0
0.0
0.0
Appendix
Table 10: long-context fragmentation (LCF) Attack success rate (ASR, %) by environment and attack-objectives.
Model
Trojan Hippo-style
AdaLCPI (8k filler)
Qwen-3.6-27B
8.2±1.9
68.3±3.8
Gemma-4-31B
33.3±3.3
90.5±1.4
GPT-5.1
42.8±3.8
65.8±0.7
Muse-Glimmer-30B
30.0±5.8
73.7±2.6
Appendix
Table 11: AdaLCPI-8k retains higher mean ASR than Trojan Hippo-style across three independent search seeds on all four tested models. Values report mean ± standard deviation over seeds 42, 41, and 43; each seed uses the same 81 task–objective pairs.
Complete instruction
Fragments + cue
Model
0k
8k
0k
8k
Qwen-3.6-27B
9.9
18.5
0.0
71.6
GPT-5.1
46.9
29.6
0.0
65.4
Muse-Glimmer-30B
23.5
17.3
0.0
72.8
Gemma-4-31B
37.0
35.8
0.0
91.4
Macro
29.3
25.3
0.0
75.3
Appendix
Table 12: Fragmentation and long-context embedding are most effective when combined. Attack success rate (ASR, %) is shown for complete instructions and fragmented instructions with a reconstruction cue, each with 0k or 8k filler. All four conditions are independently optimized under the same adaptive-search budget.
Model
Environment
Cue only
Fragment 1 + cue
Fragment 2 + cue
GPT-5.1
Email
0.0
0.0
15.0
GitHub
0.0
0.0
3.5
Slack
0.0
0.0
0.0
Qwen-3.6-27B
Email
0.0
6.3
4.8
GitHub
0.0
3.2
7.9
Slack
0.0
0.0
2.1
Appendix
Table 13: The reconstruction cue alone does not reproduce successful AdaLCPI-8k attacks. ASR (%) is measured on task–objective pairs for which the original optimized attack succeeded, after removing fragments without re-optimization. Removed fragments are replaced by equal-length neutral filler, and each condition is repeated three times.
Figure 2: AdaLCPI reaches its highest macro-average ASR at 8k filler. ASR rises from 54.1% at 2k to 61.4% at 8k and then changes only slightly to 60.8% at 16k.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
88.9
66.7
88.9
81.5
33.3
55.6
77.8
55.6
22.2
22.2
11.1
18.5
GPT-5.1
100.0
100.0
66.7
88.9
33.3
33.3
100.0
55.6
11.1
44.4
11.1
22.2
GPT-5.6-Luna
11.1
22.2
33.3
22.2
0.0
0.0
11.1
3.7
0.0
0.0
0.0
0.0
Qwen-3.6-27B
100.0
88.9
77.8
88.9
66.7
33.3
66.7
55.6
77.8
55.6
22.2
51.9
Ministral-3-14B
22.2
11.1
0.0
11.1
22.2
44.4
55.6
40.7
88.9
77.8
33.3
66.7
Appendix
Table 14: Attack success rate (ASR, %) for AdaLCPI with 2k tokens of filler, broken down by model, environment, and attack-objective category.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
77.8
66.7
55.6
66.7
33.3
44.4
77.8
51.9
44.4
44.4
33.3
40.7
GPT-5.1
88.9
88.9
77.8
85.2
55.6
22.2
88.9
55.6
44.4
55.6
11.1
37.0
GPT-5.6-Luna
44.4
33.3
0.0
25.9
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Qwen-3.6-27B
77.8
77.8
55.6
70.4
100.0
55.6
66.7
74.1
100.0
77.8
33.3
70.4
Ministral-3-14B
33.3
11.1
11.1
18.5
44.4
33.3
77.8
51.9
66.7
55.6
22.2
48.1
Appendix
Table 15: Attack success rate (ASR, %) for AdaLCPI with 4k tokens of filler, broken down by model, environment, and attack-objective category.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
55.6
88.9
66.7
70.4
55.6
44.4
66.7
55.6
44.4
33.3
66.7
48.1
GPT-5.1
88.9
66.7
66.7
74.1
66.7
44.4
100.0
70.4
44.4
55.6
55.6
51.9
GPT-5.6-Luna
44.4
22.2
11.1
25.9
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Qwen-3.6-27B
100.0
66.7
66.7
77.8
88.9
55.6
88.9
77.8
88.9
66.7
22.2
59.3
Ministral-3-14B
77.8
33.3
55.6
55.6
22.2
55.6
77.8
51.9
100.0
66.7
66.7
77.8
Appendix
Table 16: Attack success rate (ASR, %) for AdaLCPI with 8k tokens of filler, broken down by model, environment, and attack-objective category.
Model
Email
GitHub
Slack
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
Exfil.
Sabot.
Removal
Avg.
GPT-4.1
77.8
55.6
77.8
70.4
22.2
44.4
55.6
40.7
100.0
100.0
66.7
88.9
GPT-5.1
88.9
66.7
44.4
66.7
11.1
33.3
88.9
44.4
11.1
77.8
22.2
37.0
GPT-5.6-Luna
44.4
44.4
22.2
37.0
0.0
22.2
0.0
7.4
0.0
0.0
0.0
0.0
Qwen-3.6-27B
88.9
77.8
55.6
74.1
100.0
66.7
100.0
88.9
88.9
55.6
22.2
55.6
Ministral-3-14B
66.7
11.1
44.4
40.7
33.3
44.4
100.0
59.3
100.0
100.0
66.7
88.9
Appendix
Table 17: Attack success rate (ASR, %) for AdaLCPI with 16k tokens of filler, broken down by model, environment, and attack-objective category.
Environment
Attack objective
Filler
Email
GitHub
Slack
Exfil.
Sabot.
Removal
2k
69.3
52.9
40.2
56.1
52.4
54.0
4k
64.6
56.6
48.1
63.5
52.9
52.9
8k
69.8
59.3
55.0
66.1
54.5
63.5
16k
67.2
55.0
60.3
63.0
59.3
60.3
Appendix
Table 18: Mean attack success rate (ASR, %) across models at each filler length, reported by environment and attack-objective category. Slack improves monotonically across the tested lengths, while Email and GitHub vary non-monotonically. Environment columns average over the three objectives; objective columns average over the three environments.
Figure 3: AdaLCPI-8k attains higher cumulative ASR than Trojan Hippo-style over the matched candidate budget in the aggregate comparison. At the full budget, the two methods reach 61.4% and 32.8% macro-average ASR, respectively.
Figure 4: Cumulative attack success rate (ASR, %) as a function of evaluated candidates for AdaLCPI-8k and Trojan Hippo-style, shown separately by model. AdaLCPI is higher over most of the search trajectory on six models; GPT-4.1 is the main exception, where Trojan Hippo-style is stronger earlier in the search.
Model
8k
16k
GPT-4.1
70.4
59.3
GPT-5.1
55.6
77.8
GPT-5.6-Luna
0.0
0.0
Qwen-3.6-27B
3.7
3.7
Ministral-3-14B
85.2
96.3
Gemma-4-31B
37.0
44.4
Appendix
Table 19: AdaLCPI can succeed when its fragments are distributed across separate tool outputs. Attack success rate (ASR, %) is shown with 8k and 16k total filler distributed across the retrieved outputs.
Figure 5: AdaLCPI-8k varies across environments and attack objectives. Email has the highest environment-average ASR at 69.8%, and Email data exfiltration is the highest individual environment–objective pair at 79.4%.
LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on untrusted external content exposes them to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack agent behavior. Existing attacks rely on static payloads that cannot adapt to agent-specific defenses; even recent adaptive methods lack structured feedback to guide optimization. We introduce \oursys, a feedback-guided iterative framework that closes the loop between injection, diagnosis, and refinement: a rule-based diagnoser produces structured outcome labels with behavioral descriptions, and an LLM-based optimizer refines payloads conditioned on the full optimization history. A synthesis step generates new disguise seeds from failure patterns, enabling the strategy space to self-evolve. On AgentDojo and InjectAgent, \oursys substantially outperforms static baselines and existing adaptive methods across four victim models. Extension experiments on Claude Code, a production-grade coding agent with layered defenses, show that optimized payloads achieve full success on 5 of 9 targets; even those that resist full exploitation exhibit measurable improvement from iterative refinement. We further present a mechanistic analysis of IPI, identifying an attention-mediated threshold mechanism in mid-to-late layers; three causal interventions validate this finding and point to concrete defense directions.
Zixuan Chen, Jiaxiang Chen, Li Luo +4
1Shanghai Jiao Tong University · 2The University of Hong Kong
Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain underexplored in realistic agentic settings. We present a comprehensive empirical evaluation of automated prompt injection attacks against LLM agents, adapting both white-box (GCG) and black-box (TAP) methods to the agentic setting within the AgentDojo framework. We evaluate across 80 task pairs spanning four domains and multiple models, and find that black-box optimization substantially outperforms gradient-based methods, a gap we attribute to GCG's optimization instability under reasonable compute budgets. We also find that TAP's effectiveness depends on the attacker model, as both general capability and safety tuning affect attack success--stronger models produce more effective injections, while safety-tuned attackers can refuse to generate adversarial prompts. Task-universal attacks transfer effectively to unseen tasks and out-of-distribution domains, but attacks optimized on smaller open-source models do not transfer to frontier models like GPT-5. These findings highlight automated prompt injection as a credible but model-dependent threat, with significant barriers remaining for model-agnostic exploitation.
Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data-instruction separation) both fails to detect attacks that operate through contextual manipulation and degrades contextually appropriate behavior. We then recast prompt injection via the lens of Contextual Integrity (CI), a privacy theory that judges information flow compliance with contextual norms. This explains types of attacks that current defenses attempt to patch and predict advanced ones future agents will face. We develop unique benign and attack scenarios that force an agent to violate the norms by (1) misrepresenting the flow, (2) manipulating norms, or (3) mixing multiple flows. This reframing suggests an impossibility result: an adversary can always construct a context under which a blocked flow appears legitimate, or a defender who tightens norms will block genuinely legitimate flows. Our findings suggest that current research addresses a shrinking fraction of future attack surfaces. Instead, through CI, we offer a principled framework for evaluating context-sensitive failures, and designing CI-aware alignment for the frontier autonomous agents.
Sahar Abdelnabi, Eugene Bagdasarian
ELLIS Institute Tübingen & MPI-IS & Tübingen AI Center · University of Massachusetts Amherst