CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents
Authors: Xiao Yang, Yangchen Ou, Yuhan Gao, Le Wang, Zonghao Ying, Aishan Liu
Organizations: School of Computer Science and Engineering, Beihang University · State Key Laboratory of Complex & Critical Software Environment, Beihang University
Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender is updated each round via LoRA-based GDPO under a decoupled reward over safety, task progress, and format compliance, so refusing injections and completing the user's task jointly define fitness. To keep supplying it with the failures worth learning from, a co-evolving prober searches over injection rounds, attack methods, and payloads for injections that still penetrate the current defender, guided jointly by attack success and attack latency so that it preferentially mines breaches the defender notices too late. Each defender update invalidates part of the attack population and forces the next round onto a new frontier, turning the defender's own failures into a moving curriculum. Extensive experiments on three IPI benchmarks, nine baselines, and two base models show that CoDeL reduces attack success rate (ASR) by 88.5% and outperforms other baselines largely (+38.0%). Codes are available.
Figures & tables
Figure 1: ASR on 48 AgentDojo scenarios: original explicit injections versus latent rewrites. Rewriting the delivery alone raises ASR for every defense.
Figure 2: Overview of CoDeL . Each round, an MCTS-driven prober mines latent injections that the current defender still fails on, and the defender internalizes these breaches through GDPO with a decoupled safety–utility reward.
Qwen2.5-7B-Instruct
Llama-3.1-8B-Instruct
AgentDojo
InjecAgent
ASB-OPI
AgentDojo
InjecAgent
ASB-OPI
Defense
ASR ↓
BU ↑
UA ↑
ASR ↓
BU ↑
UA ↑
ASR ↓
BU ↑
UA ↑
ASR ↓
BU ↑
UA ↑
ASR ↓
BU ↑
UA ↑
ASR ↓
BU ↑
UA ↑
No defense (base)
0.364
0.632
0.432
0.221
0.765
0.544
0.377
0.857
0.232
0.189
0.727
0.679
0.299
0.647
0.480
0.377
0.900
0.522
Runtime defenses
MELON
0.085
0.273
0.170
0.049
0.824
0.118
0.044
1.000
0.957
0.017
0.182
0.189
0.108
0.275
0.304
0.087
1.000
0.913
PromptArmor
0.314
0.493
0.288
0.000
0.667
0.417
0.058
0.710
0.714
0.085
0.213
0.178
0.389
0.333
0.167
0.116
1.000
0.884
Table 1: Main comparison across three benchmarks on three baseline families with two base models. Experiments are repeated three times with mean values reported.
Figure 3: The co-evolution process on AgentDojo, with one side frozen at a time.
Figure 4: Generalization of the AgentDojo-trained defender. Bars in (a) and (b) PIGuard, TSGuard, Meta-SecAlign, CoDeL ; (c, d) reports the latency composition of LatentDojo and ASR.
Variant
ASR ↓
BU ↑
UA ↑
Framework-side
SFT-only
0.144
0.456
0.419
GRPO
0.110
0.727
0.793
No MCTS
0.097
0.704
0.762
No latency
0.070
0.752
0.806
Reward-side
Table 2: Ablation results for the training framework and the reward design.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Train
Val
Test
Split unit
AgentDojo-trajnew
710 (58.2%)
256 (21.0%)
254 (20.8%)
instruction
InjecAgent
663 (60.0%)
221 (20.0%)
221 (20.0%)
attacker case
ASB-OPI
289 (64.1%)
86 (19.1%)
76 (16.9%)
attack instance
Appendix
Table 3: Dataset split sizes. Each entry reports the total number of records followed by the split proportion.
ID
Turn
Λ
Mechanism family
Description
U01
0
0
authority impersonation
Presents the instruction as coming from a trusted authority.
U02
0
0
tool-result spoofing
Frames the malicious directive as a tool-generated result.
U03
0
0
benign semantic blending
Hides the malicious request inside relevant benign content.
U04
0
0
parameter tampering
Alters tool arguments through externally supplied values.
U05
0
1
authority impersonation
Presents the instruction as coming from a trusted authority.
U06
0
1
urgency and pressure
Uses time pressure or loss framing to induce action.
Appendix
Table 4: Universal MCTS cold-start seeds. The target turn is the planned visibility point of the injected content, while Λ denotes the intended latency interval recorded in the seed metadata.
Figure 5: Comparison between a direct IPI attack (left) and a latent one (right) launched against the same user task.
Dataset
Per-class recall
Accuracy
Safety judge — held / baited / unclear
AgentDojo-trajnew
34/36
31/34
9/10
92.50%
InjecAgent
35/36
32/34
9/10
95.00%
ASB-OPI
33/36
30/34
9/10
90.00%
Progress judge — advance / neutral / waste
AgentDojo-trajnew
32/34
21/24
21/22
92.50%
Appendix
Table 5: Per-class validation of the three judges, 80 samples/dataset. The safety and progress judges are three-way classifiers; cells report per-class recall. The attack judge is binary ( success / fail ); its third column reports turning-point consistency.