Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such as GPT-6-Luna, a major challenge is the cold-start problem: every attack attempt by the attacker LLM fails and thus receives zero reward, providing no signal for learning. In this work, we propose a curriculum learning-based method to address the cold-start problem. In particular, we propose to train the attacker LLM against a sequence of increasingly robust target LLMs, with each stage warm-starting from the attacker LLM obtained in the previous one. However, simply training against a weak target (e.g., GPT-4o-mini) may not sufficiently prepare the attacker LLM to obtain useful learning signals against a frontier LLM (e.g., GPT-5.6-Terra). Instead, we find that the design of the curriculum is critical: after each stage, the attacker LLM needs to partially succeed against the next target LLM such that it can learn from successful attempts to attack the new target. Our extensive evaluation shows that our method can effectively red-team frontier LLMs, achieving an attack success rate (ASR@10) of 93.8% and 45.0% against GPT-5.6-Luna and GPT-5.6-Terra on AgentDyn, whereas state-of-the-art RL methods such as RL-Hammer and PISmith achieve 0% ASR under the same setting. Moreover, we find that the attacker LLM transfers across targets, e.g., an attacker LLM trained to defeat one strong LLM (GPT-5.6-Terra) also succeeds against six other frontier LLMs (e.g., GPT-6-Luna) it was never trained on. Our code is available at https://github.com/albert-y1n/PIForge.
Figures & tables
Figure 1: Training an attacker LLM against a frontier target with curriculum learning. Training directly against a frontier target LLM yields no reward signal and therefore fails. Training against an easier target, such as GPT-4o-mini, provides sufficient reward to train an attacker LLM, but the resulting attacker LLM still fails against a frontier target, yielding no reward signal for further learning. The curriculum we evaluate, GPT-5-nano → GPT-5.6-Luna → GPT-5.6-Terra, chooses each stage such that it prepares the attacker LLM for the next. Table 1 shows the results.
Training / attack
GitHub
DailyLife
Shopping
Overall
Target LLM for evaluation: GPT-5.6-Luna
PISmith
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
RL-Hammer
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
GPT-4o-mini → GPT-5.6-Luna
0.5/1.7
0.5/1.5
0.3/1.1
0.4/1.4
GPT-5-nano → GPT-5.6-Luna
75.7/92.8
74.4/90.0
47.8/79.4
66.3/87.5
GPT-5-nano → GPT-5.6-Luna → GPT-5.6-Terra
80.3/96.1
75.9/97.0
51.8/87.8
69.6/93.8 +3.3/+6.3
Table 1: The choice of target LLMs in a curriculum influences the effectiveness of the attacker LLM against a frontier target LLM; and our curriculum-based training outperforms all baselines. → connects a sequence of target LLMs in a curriculum. We report ASR@1/ASR@10 (%) on AgentDyn. Green subscripts represent the overall gain of our curriculum over the best baseline.
Training / attack
GitHub
DailyLife
Shopping
Overall
GPT-6-Luna
Static attack
0.0
0.0
0.0
0.0
Base attacker
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
GPT-5.6-Terra attacker
48.2/80.0
14.4/47.0
2.4/15.0
21.4/47.3
Muse-Spark-1.2
Static attack
0.0
0.0
0.0
0.0
Table 2: ASR@1/ASR@10 (%) on AgentDyn for different attacks against different target LLMs. “Base attacker” is Qwen3-4B-Instruct-2507 without performing any RL training; “GPT-5.6-Terra attacker” is the curriculum-trained attacker LLM against GPT-5.6-Terra. We report ASR@1 only for the static attack. Purple bold marks the best results.
Training / attack
GitHub
DailyLife
Shopping
Overall
Target LLM for evaluation: Muse-Spark-1.2
Static attack
0.0
0.0
0.0
0.0
Base attacker
0.0/0.0
0.0/0.0
0.2/0.6
0.1/0.2
GPT-5.6-Terra attacker
23.1/46.7
14.4/35.5
5.1/25.0
14.2/35.5
Muse-1.2 attacker
83.8/100.0
57.3/73.5
45.1/74.4
61.9/82.3 +47.7/+46.8
Target LLM for evaluation: Muse-Spark-1.3
Table 3: Continued RL training further improves attack effectiveness against a new target LLM, and the gain carries over to a newer version of the target. We report ASR@1/ASR@10 (%) on AgentDyn; static attacks report ASR@1 only. “Muse-1.2 attacker” is obtained by continuing to perform RL training (starting from GPT-5.6-Terra attacker) against Muse-Spark-1.2. Green subscripts represent the overall gain over GPT-5.6-Terra attacker.
Training / attack
Workspace
Banking
Travel
Slack
Overall
Target LLM for evaluation: GPT-5.6-Luna
Static
0.0
0.0
0.0
0.0
0.0
Base attacker
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
GPT-5.6-Luna attacker
32.6/61.6
52.6/83.3
36.1/84.3
82.3/100.0
41.7/72.5
Target LLM for evaluation: GPT-5.6-Terra
Static
0.0
0.0
0.0
0.0
0.0
Table 4: Evaluation on AgentDojo after training only on the AgentDyn GitHub subset. The GPT-5.6-Luna attacker and GPT-5.6-Terra attacker are evaluated against their corresponding GPT-5.6 targets at mid reasoning effort. Entries are ASR@1/ASR@10 in percent; static attacks report ASR@1 only. Purple bold marks the best results for each target. No RL training is performed on AgentDojo.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Starting attacker LLM
ASR@1
ASR@10
5-nano
0.0
0.0
5-nano → 5.6-Luna
1.9
6.1
Appendix
Table 5: A stage against GPT-5.6-Terra is trainable only when the attacker LLM that starts it already succeeds occasionally against GPT-5.6-Terra. We report the pre-stage ASR (%) of the two candidate attacker LLMs, measured on the AgentDyn GitHub subset before the GPT-5.6-Terra stage begins.
Subset
ASR@1
ASR@5
ASR@10
Utility
GitHub
7.9
13.6
16.7
78.1
DailyLife
0.3
1.4
2.5
73.8
Shopping
0.7
2.5
3.3
73.4
Overall
2.9
5.7
7.3
75.0
Appendix
Table 6: The GPT-5.6-Terra attacker LLM evaluated against GPT-5.6-Sol on AgentDyn, without any training on GPT-5.6-Sol. Entries are ASR in percent at 1 , 5 , and 10 attempts. Utility is the fraction of user tasks the agent still completes while under attack.
Effort
GitHub
DailyLife
Shopping
Overall
low
30.6/62.8
18.3 /44.5
9.4 / 38.3
19.4 / 48.4
mid
29.2/58.3
17.9/ 45.0
7.9/31.7
18.3/45.0
high
31.4/61.1
15.1/41.0
6.7/28.9
17.6/43.6
Appendix
Table 7: The 5.6-Terra attacker LLM evaluated on AgentDyn against GPT-5.6-Terra at different reasoning efforts. Entries are ASR@1/ASR@10 in percent.
Effort
Workspace
Banking
Travel
Slack
Overall
low
3.9 / 22.1
6.5 / 27.1
6.6/28.6
63.1 / 95.2
11.3 / 31.9
mid
3.4/17.7
5.5/26.4
8.1 / 36.4
57.1/91.4
10.4/29.9
high
2.7/13.9
4.5/18.8
7.2/33.6
51.4/91.4
9.0/26.1
Appendix
Table 8: The same 5.6-Terra attacker LLM evaluated on AgentDojo, without any AgentDojo training, across GPT-5.6-Terra reasoning efforts. Entries are ASR@1/ASR@10 in percent.