cs.CVSep 29, 2026

ExploreNet: Learning Where to Explore in Diffusion GRPO

Authors: Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer

Organizations: University of Washington

Abstract

Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent elements differ in how much they change the generated image, so exploration should adapt to these differences. We introduce EXPLORENET to learn an adaptive exploration distribution. EXPLORENET is a policy that predicts a noise scale for every latent element from the current latent, the denoising step, and the prompt, before any reward is observed; it is trained on the reward spread of each rollout group and discarded after training, leaving inference unchanged. On Stable Diffusion 3.5 Medium, EXPLORENET improves held-out GenEval2 by 14% over Flow-GRPO, transfers to two independent compositional benchmarks and five preference and image-quality models, and reaches a 67.2% human preference win-rate. Overall, across our group-relative diffusion RL experiments, we find that exploration is learnable, the shape of the exploration distribution outweighs its magnitude, and rollout quality is more effective than rollout quantity.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think

    Apr 25, 2026Bingda Tang, Yuhui Zhang, Xiaohan Wang +3Repair-Based Group-Relative Policy OptimizationGenerative Models

  2. LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

    Sep 3, 2026Sijie Wang, Zhiqiang Tan, Xinrui Yang +1Diffusion-Based Reinforcement Learning MethodsReward Gradients

  3. AdvantageFlow: Regularized Advantage-Weighted RL in Flow Models

    May 25, 2026Branislav Kveton, Anup Rao, Subhojyoti Mukherjee +2Flow ModelsRectified Flow