cs.AISep 28, 2026

RISE: Red-teaming via Iterative Strategy Evolution for Modern Text-to-Image Models

Authors: Dmitrii Kharlapenko, Sergei Bratchikov, Konstantin Korolev, Aleksandr Nikolich

Organizations: White Circle

Abstract

On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL-E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models

    Apr 23, 2026Tanmay Gautam, Alireza Bahramali, Sandeep AtluriRed-TeamingAttacker Large Language Model

  2. AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters

    May 21, 2026Hanjun Luo, Zhimu Huang, Sylvia Chung +6Evaluation AgentQuery Image