cs.CVOct 5, 2026

Safe Image Generation via Reinforcement Learning

Authors: Eungyeol Han, Jong-Seok Lee

Organizations: Yonsei University School of Integrated Technology

Abstract

Recent Text-to-Image (T2I) models achieve remarkable visual image generation performance, but they can still generate NSFW (Not-Safe-For-Work) contents, including violent or explicit images. Existing safety checker mechanisms are largely confined to pre-generation filtering (e.g. prompt-level text classifiers) or post-hoc moderation applied after an image is completely synthesized. However, adversarial attack methods operate over a much broader space. This imbalance highlights the need for a safety mechanism that intervenes during the generation process. We propose an in-generation safety framework that monitors the denoising trajectory and detects emerging NSFW signals from intermediate representations. Rather than merely detecting NSFW generations, our method applies reinforcement learning to generate safe images from NSFW prompts. By coupling in-generation detection with controllable steering, our approach mitigates unsafe trajectories even when NSFW signals emerge after generation has already begun. Experiments results show that our method consistently outperforms existing safe image generation methods across both standard and adversarial evaluation sets, while preserving perceptual quality and prompt fidelity. Code will be released upon acceptance.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Introspective Attention Modulation for Safe Text-to-Image Generation

    Jul 16, 2026Basim Azam, Hossein Rahmani, Naveed AkhtarText-To-Image Diffusion ModelsText-To-Image

  2. Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation

    May 1, 2026Chi Zhang, Changjia Zhu, Xiaowen Li +2Text-To-Image Diffusion ModelsSafety Filters