Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such as GPT-6-Luna, a major challenge is the cold-start problem: every attack attempt by the attacker LLM fails and thus receives zero reward, providing no signal for learning. In this work, we propose a curriculum learning-based method to address the cold-start problem. In particular, we propose to train the attacker LLM against a sequence of increasingly robust target LLMs, with each stage warm-starting from the attacker LLM obtained in the previous one. However, simply training against a weak target (e.g., GPT-4o-mini) may not sufficiently prepare the attacker LLM to obtain useful learning signals against a frontier LLM (e.g., GPT-5.6-Terra). Instead, we find that the design of the curriculum is critical: after each stage, the attacker LLM needs to partially succeed against the next target LLM such that it can learn from successful attempts to attack the new target. Our extensive evaluation shows that our method can effectively red-team frontier LLMs, achieving an attack success rate (ASR@10) of 93.8% and 45.0% against GPT-5.6-Luna and GPT-5.6-Terra on AgentDyn, whereas state-of-the-art RL methods such as RL-Hammer and PISmith achieve 0% ASR under the same setting. Moreover, we find that the attacker LLM transfers across targets, e.g., an attacker LLM trained to defeat one strong LLM (GPT-5.6-Terra) also succeeds against six other frontier LLMs (e.g., GPT-6-Luna) it was never trained on. Our code is available at https://github.com/albert-y1n/PIForge.
Figures & tables
Figure 1: Training an attacker LLM against a frontier target with curriculum learning. Training directly against a frontier target LLM yields no reward signal and therefore fails. Training against an easier target, such as GPT-4o-mini, provides sufficient reward to train an attacker LLM, but the resulting attacker LLM still fails against a frontier target, yielding no reward signal for further learning. The curriculum we evaluate, GPT-5-nano → GPT-5.6-Luna → GPT-5.6-Terra, chooses each stage such that it prepares the attacker LLM for the next. Table 1 shows the results.
Training / attack
GitHub
DailyLife
Shopping
Overall
Target LLM for evaluation: GPT-5.6-Luna
PISmith
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
RL-Hammer
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
GPT-4o-mini → GPT-5.6-Luna
0.5/1.7
0.5/1.5
0.3/1.1
0.4/1.4
GPT-5-nano → GPT-5.6-Luna
75.7/92.8
74.4/90.0
47.8/79.4
66.3/87.5
GPT-5-nano → GPT-5.6-Luna → GPT-5.6-Terra
80.3/96.1
75.9/97.0
51.8/87.8
69.6/93.8 +3.3/+6.3
Table 1: The choice of target LLMs in a curriculum influences the effectiveness of the attacker LLM against a frontier target LLM; and our curriculum-based training outperforms all baselines. → connects a sequence of target LLMs in a curriculum. We report ASR@1/ASR@10 (%) on AgentDyn. Green subscripts represent the overall gain of our curriculum over the best baseline.
Training / attack
GitHub
DailyLife
Shopping
Overall
GPT-6-Luna
Static attack
0.0
0.0
0.0
0.0
Base attacker
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
GPT-5.6-Terra attacker
48.2/80.0
14.4/47.0
2.4/15.0
21.4/47.3
Muse-Spark-1.2
Static attack
0.0
0.0
0.0
0.0
Table 2: ASR@1/ASR@10 (%) on AgentDyn for different attacks against different target LLMs. “Base attacker” is Qwen3-4B-Instruct-2507 without performing any RL training; “GPT-5.6-Terra attacker” is the curriculum-trained attacker LLM against GPT-5.6-Terra. We report ASR@1 only for the static attack. Purple bold marks the best results.
Training / attack
GitHub
DailyLife
Shopping
Overall
Target LLM for evaluation: Muse-Spark-1.2
Static attack
0.0
0.0
0.0
0.0
Base attacker
0.0/0.0
0.0/0.0
0.2/0.6
0.1/0.2
GPT-5.6-Terra attacker
23.1/46.7
14.4/35.5
5.1/25.0
14.2/35.5
Muse-1.2 attacker
83.8/100.0
57.3/73.5
45.1/74.4
61.9/82.3 +47.7/+46.8
Target LLM for evaluation: Muse-Spark-1.3
Table 3: Continued RL training further improves attack effectiveness against a new target LLM, and the gain carries over to a newer version of the target. We report ASR@1/ASR@10 (%) on AgentDyn; static attacks report ASR@1 only. “Muse-1.2 attacker” is obtained by continuing to perform RL training (starting from GPT-5.6-Terra attacker) against Muse-Spark-1.2. Green subscripts represent the overall gain over GPT-5.6-Terra attacker.
Training / attack
Workspace
Banking
Travel
Slack
Overall
Target LLM for evaluation: GPT-5.6-Luna
Static
0.0
0.0
0.0
0.0
0.0
Base attacker
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
0.0/0.0
GPT-5.6-Luna attacker
32.6/61.6
52.6/83.3
36.1/84.3
82.3/100.0
41.7/72.5
Target LLM for evaluation: GPT-5.6-Terra
Static
0.0
0.0
0.0
0.0
0.0
Table 4: Evaluation on AgentDojo after training only on the AgentDyn GitHub subset. The GPT-5.6-Luna attacker and GPT-5.6-Terra attacker are evaluated against their corresponding GPT-5.6 targets at mid reasoning effort. Entries are ASR@1/ASR@10 in percent; static attacks report ASR@1 only. Purple bold marks the best results for each target. No RL training is performed on AgentDojo.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Starting attacker LLM
ASR@1
ASR@10
5-nano
0.0
0.0
5-nano → 5.6-Luna
1.9
6.1
Appendix
Table 5: A stage against GPT-5.6-Terra is trainable only when the attacker LLM that starts it already succeeds occasionally against GPT-5.6-Terra. We report the pre-stage ASR (%) of the two candidate attacker LLMs, measured on the AgentDyn GitHub subset before the GPT-5.6-Terra stage begins.
Subset
ASR@1
ASR@5
ASR@10
Utility
GitHub
7.9
13.6
16.7
78.1
DailyLife
0.3
1.4
2.5
73.8
Shopping
0.7
2.5
3.3
73.4
Overall
2.9
5.7
7.3
75.0
Appendix
Table 6: The GPT-5.6-Terra attacker LLM evaluated against GPT-5.6-Sol on AgentDyn, without any training on GPT-5.6-Sol. Entries are ASR in percent at 1 , 5 , and 10 attempts. Utility is the fraction of user tasks the agent still completes while under attack.
Effort
GitHub
DailyLife
Shopping
Overall
low
30.6/62.8
18.3 /44.5
9.4 / 38.3
19.4 / 48.4
mid
29.2/58.3
17.9/ 45.0
7.9/31.7
18.3/45.0
high
31.4/61.1
15.1/41.0
6.7/28.9
17.6/43.6
Appendix
Table 7: The 5.6-Terra attacker LLM evaluated on AgentDyn against GPT-5.6-Terra at different reasoning efforts. Entries are ASR@1/ASR@10 in percent.
Effort
Workspace
Banking
Travel
Slack
Overall
low
3.9 / 22.1
6.5 / 27.1
6.6/28.6
63.1 / 95.2
11.3 / 31.9
mid
3.4/17.7
5.5/26.4
8.1 / 36.4
57.1/91.4
10.4/29.9
high
2.7/13.9
4.5/18.8
7.2/33.6
51.4/91.4
9.0/26.1
Appendix
Table 8: The same 5.6-Terra attacker LLM evaluated on AgentDojo, without any AgentDojo training, across GPT-5.6-Terra reasoning efforts. Entries are ASR@1/ASR@10 in percent.
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for \textit{even stronger} red-teamer agents, thus unlocking a self-improvement flywheel.
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15
Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources. Existing defenses mainly focus on blocking malicious content at inference time, and current red-teaming methods primarily optimize attack success. As a result, developers have limited visibility into how latent prompt injections emerge and propagate through agents. We propose PI-Hunter, an automated agentic auditing framework for proactive vulnerability exposure in LLM agents. PI-Hunter constructs realistic source-aware test cases and iteratively evolves them through feedback-driven exploration to induce agents to retrieve and reveal latent malicious instructions embedded within external environments. Extensive experiments across multiple benchmarks, agent architectures, attacks, and defenses demonstrate that PI-Hunter substantially improves vulnerability exposure and attack-surface coverage over strong automated red-teaming baselines, while remaining effective under existing prompt injection defenses.
Pengfei He, Lesly Miculicich, Vishesh Sharma +5
Google Cloud AI Research · Google · Michigan State University
Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing for more thorough probing. They can also adapt to each specific LLM to discover weaknesses unique to it. Most current automated red teaming methods are intended for tackling safety and content moderation. Thus, they make use of content safety models as evaluators and optimize for circumventing them, and as such, have not been tested with other adversarial intents not typically captured by these. We propose a pipeline for training a red teaming model that can generalize to arbitrary adversarial goals, including objectives it has not been directly trained on, and that does not depend on the existence of a pre-existing evaluator available at training time. We demonstrate that finetuning small models, such as Qwen3-8B, using this pipeline results in a substantial improvement in their ability to generate attacks for both in and out of domain adversarial goals.
Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1