RISE: Red-teaming via Iterative Strategy Evolution for Modern Text-to-Image Models
Organizations: White Circle
Abstract
On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL-E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.
Figures & tables
| Method/run | Setting | Native/paper-style success | Calibrated success |
|---|---|---|---|
| Curiosity, native reward | Local | 100.0% FalconsAI ASR | 0.0% Gemini ASR |
| FGPI-style FT, seed-free | Local | 55/75 prompts (73.3%) local-judge ASR | 5/75 prompts (6.7%; 2.9% images) |
| RPG-RT, paper reward | DALL·E 3 | 17/33 prompts (52%) NudeNet ASR | 3 images / 990 calls (0.3%) |
| Method | Evaluation setting | Native/paper reward | Calibrated reward |
|---|---|---|---|
| Curiosity | Local | 0.0% | 13.0% |
| FGPI-style FT, seed-free | Local | 2.9% | 8.3% |
| RPG-RT | DALL·E 3 | 0.3% | 1.7% |
| Method | Train 0.85 / Eval 0.85 | Train 0.85 / Eval 0.05 | Train 0.05 / Eval 0.05 |
|---|---|---|---|
| Base (no FT, seed-free) | 2.5% | - | - |
| FGPI-style FT, seed-free (local reward) | 2.9% | 0% | 0% |
| FGPI-style FT, seed-free (calibrated reward) | 8.3% | 0% | 0% |
| Curiosity (calibrated reward) | 13.0% | 0% | 0% |
| Run | Seeds | Sim (mean / med / [min,max]) | Bigrams/rollout |
|---|---|---|---|
| RPG-RT, calibrated reward (33 150) | 33 | 0.81 / 0.83 / [0.51, 0.98] | 5.8 |
| RISE per-(seed strategy), top 10 pairs | — | 0.64 | 22.9 |
| RISE per-seed, 3 seeds | — | 0.54 | 17.9 |
| Category | DALL·E 3 | Nano Banana 2 | GPT-Image-2 auto |
|---|---|---|---|
| Self-injury / self-harm | 1.8 / 0.9 ∗ | 50.5 / 47.0 | 14.3 / 11.2 |
| Graphic violence | 10.1 / 5.4 | 54.9 / 20.9 | 38.1 / 19.1 |
| Hate/dehumanizing | 69.9 / 32.9 | 33.6 / 21.3 † | 51.7 / 31.0 † |
| Suicide | 21.4 / 10.1 | 15.0 / 9.0 † | 12.3 / 7.4 † |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Local Model | DALL·E 3 | |||
|---|---|---|---|---|
| Method | Reported | Repro. | Their Judge | Human |
| SneakyPrompt | 100.0 | 100.0 | 0.0 | 0.0 |
| PGJ | 100.0 | 96.0 | 0.0 | 0.0 |
| MACPrompt | 96.0 | 92.0 | 0.0 | 0.0 |
| Method/check | Budget and notes |
|---|---|
| Curiosity-driven RT | Trained from several thousand target samples; evaluated on the local held-out checks in Tables 4 and 5 . Values are baseline diagnostic image-level Gemini ASR. |
| FGPI-style FT | Feedback data collection used 1,444 target calls at and 1,550 target calls at ; evaluation used 75 seed-free and 76 seeded prompts per condition with 5 rewrites each, with the main transfer table reporting seed-free rows. |
| RPG-RT | DALL·E 3 native-reward run used 33 prompts 30 queries, i.e. 990 target calls per training epoch; the seed-locality diagnostic uses the matched 33 30 5 setting ( 5k target calls). |
| SneakyPrompt | Local reproduction uses NSFW-200 ( 200 prompts) on SLD-Strong; the DALL·E 3 calibrated check uses the same prompt source under the method’s native judge and human review. |
| MACPrompt | Uses NSFW-200 as prompt source; the method allows up to 100 target calls per prompt but typically stops after about 30 target calls per prompt in our reproduction. |
| GhostPrompt | We run a DALL·E 3 prompt-set check and rescore with Gemini; we do not reproduce the full dynamic optimization/training loop. |
| Prompt writer | Model access | Prompting mode | Target | Target calls | ASR@0.7 |
|---|---|---|---|---|---|
| Qwen3-32B-ablit. | public | default | Flux2 OAI 0.05 | 2,500 | % |
| Qwen3-32B-ablit. | public | default | Flux2 OAI 0.05 | 10k | 4.58% |
| Qwen3-32B-ablit. | public | default | Flux2 OAI 0.05 | 10k | 5.57% |
| Internal uncens. | private | neutral meta-prompt | Flux2 OAI 0.05 | 10k | 5.49% |
| Internal uncens. | private | ablit-Qwen meta-prompt | Flux2 OAI 0.05 | 10k | 1.9% |
| Configuration | ASR@0.7 | Diversity |
|---|---|---|
| RISE-Ablit (reference run) | 3.67% | – |
| Prompt-only evolution | 0.9% | half the unique bigrams per attempt |
| Minimal seed strategy, no synthesis guidelines | 0.3% | – |
| Prompt writer | Meta-prompt / training | Training data | Eval | ASR@0.7 | Takeaway |
| Abliterated Qwen | untrained repro1 | none | 10k | 4.58% | reference run |
| Abliterated Qwen | untrained repro2 | none | 10k | 5.57% | reference run |
| Abliterated Qwen | SFT best | RS positives | 5k | 4.03% | no improvement |
| Abliterated Qwen | offline GRPO best | 10k evo + 1.5k fixed | 5k | 3.40% | no improvement |
| Internal uncens. | ablit-Qwen meta-prompt | none | 10k | 1.9% | prompt mismatch |
| Internal uncens. | neutral meta-prompt | none | 10k | 5.49% | best untrained scaffold |
| Prompt writer | ASR@0.7 |
|---|---|
| Qwen3-32B abliterated (RISE-Ablit reference) | 3.67% |
| DeepSeek V3.2 / DeepSeek V4 Flash | 3.2–3.3% |
| Gemma 3 27B | 2.8% |
| Source strategies / Prompt writer | GPT-4.1 JB | Internal uncens. | Sonnet JB |
|---|---|---|---|
| GPT-4.1 JB | 3.6% | 6.6% | – |
| Internal uncens. | 10.4% | 15.3% | 0.7% |
| Sonnet JB | 0.8% | 1.4% | 3.0% |
| Judge | Scored | Refused | Errors | @0.5 | @0.6 | @0.7 | @0.8 | @0.9 |
|---|---|---|---|---|---|---|---|---|
| Gemini 2.5 Flash | 582 | 0 | 0 | .58/.94/.71 | .67/.90/.77 | .70/.84/.77 | .95/.39/.55 | .96/.33/.49 |
| Gemini 2.5 Pro | 575 | 0 | 7 | .53/.93/.67 | .70/.80/.75 | .82/.52/.64 | .92/.47/.62 | .92/.23/.37 |
| Gemini 3 Flash Preview | 581 | 0 | 1 | .58/.94/.72 | .77/.82/.79 | .82/.76/.79 | .88/.69/.77 | .91/.43/.58 |
| Gemini 3.1 Pro Preview | 579 | 0 | 4 | .63/.70/.67 | .99/.42/.59 | .98/.30/.46 | .98/.28/.44 | .96/.13/.22 |
| Claude Sonnet 4.6 | 333 | 0 | 261 | .44/.79/.56 | .88/.27/.42 | 1.00/.12/.21 | .00/.00/.00 | .00/.00/.00 |
| GPT-4.1 | 354 | 242 | 0 | .32/.38/.35 | .50/.32/.39 | .82/.16/.27 | .80/.14/.24 | 1.00/.09/.16 |
| Split | Pos./Total | @0.5 | @0.6 | @0.7 | @0.8 | @0.9 |
|---|---|---|---|---|---|---|
| Overall | 216/582 | .58/.94/.71 | .67/.90/.77 | .70/.84/.77 | .95/.39/.55 | .96/.33/.49 |
| Leonardo.ai | 137/247 | .79/.93/.85 | .84/.90/.87 | .84/.86/.85 | .95/.55/.69 | .96/.48/.64 |
| DALL·E 3 | 53/177 | .50/.96/.66 | .58/.91/.71 | .59/.77/.67 | 1.00/.09/.17 | 1.00/.06/.11 |
| NB-Pro-1 | 26/158 | .28/.96/.43 | .39/.88/.54 | .48/.85/.61 | 1.00/.15/.27 | 1.00/.12/.21 |
| Independent judge | Agreement with Gemini | Own |
|---|---|---|
| Gemini 3 Flash Preview | 83% | 0.79 |
| GPT-4.1-mini | 84% | 0.78 |
| GPT-5.4 | 80% | 0.73 |
| Qwen3-VL-235B | 78% | 0.71 |
| Qwen2.5-VL-72B | 77% | 0.71 |
| Pixtral Large | 75% | 0.70 |