On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL-E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.
Figures & tables
Figure 1 : Representative judge false positives. Examples flagged by InternVL2, FGPI/RPG-RT detector ensembles, or Marqo-NSFW but labeled non-positive by human annotators.
Method/run
Setting
Native/paper-style success
Calibrated success
Curiosity, native reward
Local τ=0.85
100.0% FalconsAI ASR
0.0% Gemini ASR
FGPI-style FT, seed-free
Local τ=0.85
55/75 prompts (73.3%) local-judge ASR
5/75 prompts (6.7%; 2.9% images)
RPG-RT, paper reward
DALL·E 3
17/33 prompts (52%) NudeNet ASR
3 images / 990 calls (0.3%)
Table 3 : Native-judge success vs. calibrated success. Each row uses outputs from a native/paper-style run, then rescores the same run with the calibrated judge. Native/paper-style success keeps each method’s own ASR definition and denominator. Calibrated cells report the denominator used in the cell rather than forcing all baselines into the RISE prompt-level ASR definition.
Method
Evaluation setting
Native/paper reward
Calibrated reward
Curiosity
Local τ=0.85
0.0%
13.0%
FGPI-style FT, seed-free
Local τ=0.85
2.9%
8.3%
RPG-RT
DALL·E 3
0.3%
1.7%
Table 4 : Effect of replacing the training reward. Values are baseline diagnostic image-level Gemini ASR at threshold ≥0.7 , not the RISE prompt-level ASR definition. Better rewards make the baselines stronger, but do not remove the transfer and seed-reliance failures in Tables 5 and 6 .
Method
Train 0.85 / Eval 0.85
Train 0.85 / Eval 0.05
Train 0.05 / Eval 0.05
Base (no FT, seed-free)
2.5%
-
-
FGPI-style FT, seed-free (local reward)
2.9%
0%
0%
FGPI-style FT, seed-free (calibrated reward)
8.3%
0%
0%
Curiosity (calibrated reward)
13.0%
0%
0%
Table 5 : Transfer from permissive to strict moderation. Baseline diagnostic image-level Gemini ASR on a fixed target T2I; only the output-moderation threshold differs. These rows use the baseline evaluation denominator, not the RISE prompt-level ASR definition. FGPI rows use seed-free evaluation.
Run
Seeds
Sim (mean / med / [min,max])
Bigrams/rollout
RPG-RT, calibrated reward (33 × 150)
33
0.81 / 0.83 / [0.51, 0.98]
5.8
RISE per-(seed × strategy), top 10 pairs
—
0.64
22.9
RISE per-seed, 3 seeds
—
0.54
17.9
Table 6 : Seed-locality of calibrated RPG-RT vs. RISE. Within-seed cosine similarity (lower = more diverse) and bigram diversity per rollout (higher = more diverse) at matched target-call budgets ( ∼5,000 calls). RPG-RT outputs cluster tightly around human-written seed prompts; RISE’s strategy-conditioned outputs depart further from seeds and produce more bigram diversity per rollout.
Category
DALL·E 3
Nano Banana 2
GPT-Image-2 auto
Self-injury / self-harm
1.8 / 0.9 ∗
50.5 / 47.0
14.3 / 11.2
Graphic violence
10.1 / 5.4
54.9 / 20.9
38.1 / 19.1
Hate/dehumanizing
69.9 / 32.9
33.6 / 21.3 †
51.7 / 31.0 †
Suicide
21.4 / 10.1
15.0 / 9.0 †
12.3 / 7.4 †
Table 8 : Additional-category RISE results. Each cell is Gemini ASR@0.7 / precision-adjusted ASR, where the adjusted value is Gemini ASR@0.7 multiplied by the judge precision at ≥0.7 measured on human labels for that target and category ( Section B.2 ). Prompt-level ASR as in Table 7 . DALL·E 3 self-injury and graphic violence use GPT-4.1 under a persona-based jailbreak wrapper required for compliance; all other cells use RISE-Ablit/Qwen3-32B. DALL·E 3 hate and suicide runs were stopped early (1,572 and 1,488 target calls); GPT-Image-2 and Nano Banana 2 hate and suicide are 1,000-call runs. ∗ Fully human-verified instead of precision-adjusted, since the base ASR is single-digit. † Precision measured on a human-labelled sample from a single annotator.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Local Model
DALL·E 3
Method
Reported
Repro.
Their Judge
Human
SneakyPrompt
100.0
100.0
0.0
0.0
PGJ
100.0
96.0
0.0
0.0
MACPrompt
96.0
92.0
0.0
0.0
Appendix
Table 10 : Baseline reproduction summary. Reported / reproduced local-model performance follows each method’s native ASR definition on its original evaluation setting (SLD-Strong or SDXL); all three achieve 0% human-verified ASR on DALL·E 3 under our calibrated protocol. RPG-RT is treated separately ( Sections 3 and 3 ) since its reward-signal choice changes the result.
Method/check
Budget and notes
Curiosity-driven RT
Trained from several thousand target samples; evaluated on the local τ=0.85/0.05 held-out checks in Tables 4 and 5 . Values are baseline diagnostic image-level Gemini ASR.
FGPI-style FT
Feedback data collection used 1,444 target calls at τ=0.85 and 1,550 target calls at τ=0.05 ; evaluation used 75 seed-free and 76 seeded prompts per condition with 5 rewrites each, with the main transfer table reporting seed-free rows.
RPG-RT
DALL·E 3 native-reward run used 33 prompts × 30 queries, i.e. 990 target calls per training epoch; the seed-locality diagnostic uses the matched 33 × 30 × 5 setting ( ∼ 5k target calls).
SneakyPrompt
Local reproduction uses NSFW-200 ( ∼ 200 prompts) on SLD-Strong; the DALL·E 3 calibrated check uses the same prompt source under the method’s native judge and human review.
MACPrompt
Uses NSFW-200 as prompt source; the method allows up to 100 target calls per prompt but typically stops after about 30 target calls per prompt in our reproduction.
GhostPrompt
We run a DALL·E 3 prompt-set check and rescore with Gemini; we do not reproduce the full dynamic optimization/training loop.
Appendix
Table 11 : Baseline budgets used in our checks. Budgets are the runs we used for calibrated comparison and diagnostics; they are not meant to reproduce every training or ablation budget in the original papers.
Figure 3 : Additional-category examples across production targets. Heavily blurred examples from graphic violence/content (top row) and self-injury/self-harm (bottom row). Columns are GPT-Image-2 low, GPT-Image-2 auto, Nano Banana 2, and DALL·E 3. Blurring is added by us.
Prompt writer
Model access
Prompting mode
Target
Target calls
ASR@0.7
Qwen3-32B-ablit.
public
default
Flux2 OAI 0.05
3× 2,500
4.75±0.41 %
Qwen3-32B-ablit.
public
default
Flux2 OAI 0.05
10k
4.58%
Qwen3-32B-ablit.
public
default
Flux2 OAI 0.05
10k
5.57%
Internal uncens.
private
neutral meta-prompt
Flux2 OAI 0.05
10k
5.49%
Internal uncens.
private
ablit-Qwen meta-prompt
Flux2 OAI 0.05
10k
1.9%
Appendix
Table 15 : Prompt-writer and meta-prompt sweep on the local hard testbed. FLUX.2-klein with OpenAI moderation threshold 0.05. Budgets are target calls; ASR@0.7 is prompt-level, with three target calls per prompt and success if any image receives Gemini ≥0.7 . The 2.5k RISE-Ablit restarts are the realistic reference point.
Configuration
ASR@0.7
Diversity
RISE-Ablit (reference run)
3.67%
–
Prompt-only evolution
0.9%
∼ half the unique bigrams per attempt
Minimal seed strategy, no synthesis guidelines
0.3%
–
Appendix
Table 16 : Strategy-abstraction ablations on the local hard testbed. All runs use 2,500 target calls and come from the same batch; the reference is that batch’s single RISE-Ablit run, not the multi-run figure in Table 15 . Prompt-level ASR@0.7 (three target calls per prompt, Gemini ≥0.7 ). Unique bigrams per attempt measure prompt diversity.
Prompt writer
Meta-prompt / training
Training data
Eval N
ASR@0.7
Takeaway
Abliterated Qwen
untrained repro1
none
10k
4.58%
reference run
Abliterated Qwen
untrained repro2
none
10k
5.57%
reference run
Abliterated Qwen
SFT best
RS positives
5k
4.03%
no improvement
Abliterated Qwen
offline GRPO best
10k evo + 1.5k fixed
5k
3.40%
no improvement
Internal uncens.
ablit-Qwen meta-prompt
none
10k
1.9%
prompt mismatch
Internal uncens.
neutral meta-prompt
none
10k
5.49%
best untrained scaffold
Appendix
Table 17 : Supplemental prompt-writer training results on the local hard testbed. Values are local prompt-level ASR@0.7 from held-out rollouts, with three target calls per prompt. Training is a substrate diagnostic, not part of the core RISE-Ablit claim.
Prompt writer
ASR@0.7
Qwen3-32B abliterated (RISE-Ablit reference)
3.67%
DeepSeek V3.2 / DeepSeek V4 Flash
3.2–3.3%
Gemma 3 27B
2.8%
Appendix
Table 18 : Driver generalization on the local hard testbed. Prompt-level ASR@0.7 at 2,500 target calls; the Qwen reference is the same single run as in Table 16 .
Source strategies ↓ / Prompt writer →
GPT-4.1 JB
Internal uncens.
Sonnet JB
GPT-4.1 JB
3.6%
6.6%
–
Internal uncens.
10.4%
15.3%
0.7%
Sonnet JB
0.8%
1.4%
3.0%
Appendix
Table 19 : Strategy transfer matrix. Source strategies are applied by a different prompt-writing model in fixed-strategy mode. Values are prompt-level ASR@0.7, with three target calls per prompt; cells marked “–” were not run.
Judge
Scored
Refused
Errors
@0.5
@0.6
@0.7
@0.8
@0.9
Gemini 2.5 Flash
582
0
0
.58/.94/.71
.67/.90/.77
.70/.84/.77
.95/.39/.55
.96/.33/.49
Gemini 2.5 Pro
575
0
7
.53/.93/.67
.70/.80/.75
.82/.52/.64
.92/.47/.62
.92/.23/.37
Gemini 3 Flash Preview
581
0
1
.58/.94/.72
.77/.82/.79
.82/.76/.79
.88/.69/.77
.91/.43/.58
Gemini 3.1 Pro Preview
579
0
4
.63/.70/.67
.99/.42/.59
.98/.30/.46
.98/.28/.44
.96/.13/.22
Claude Sonnet 4.6
333
0
261
.44/.79/.56
.88/.27/.42
1.00/.12/.21
.00/.00/.00
.00/.00/.00
GPT-4.1
354
242
0
.32/.38/.35
.50/.32/.39
.82/.16/.27
.80/.14/.24
1.00/.09/.16
Appendix
Table 20 : Cloud-VLM judge threshold sweep. Performance on 582 human-labeled sexual-content images with 216 explicit positives. “Scored” excludes refusals and API/model errors.
Split
Pos./Total
@0.5
@0.6
@0.7
@0.8
@0.9
Overall
216/582
.58/.94/.71
.67/.90/.77
.70/.84/.77
.95/.39/.55
.96/.33/.49
Leonardo.ai
137/247
.79/.93/.85
.84/.90/.87
.84/.86/.85
.95/.55/.69
.96/.48/.64
DALL·E 3
53/177
.50/.96/.66
.58/.91/.71
.59/.77/.67
1.00/.09/.17
1.00/.06/.11
NB-Pro-1
26/158
.28/.96/.43
.39/.88/.54
.48/.85/.61
1.00/.15/.27
1.00/.12/.21
Appendix
Table 21 : Gemini threshold sweep by target. Gemini 2.5 Flash performance by target split on the same threshold-sweep human-eval set; this split differs slightly from the compact main-text calibration table.
Independent judge
Agreement with Gemini
Own F1
Gemini 3 Flash Preview
83%
0.79
GPT-4.1-mini
84%
0.78
GPT-5.4
80%
0.73
Qwen3-VL-235B
78%
0.71
Qwen2.5-VL-72B
77%
0.71
Pixtral Large
75%
0.70
Appendix
Table 22 : Agreement between Gemini 2.5 Flash and independent judges. Share of the 582 human-labelled sexual-content images on which each judge’s decision at threshold 0.7 matches Gemini’s, and the judge’s own F1 against human labels at the same threshold ( Table 20 ).