Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent elements differ in how much they change the generated image, so exploration should adapt to these differences. We introduce EXPLORENET to learn an adaptive exploration distribution. EXPLORENET is a policy that predicts a noise scale for every latent element from the current latent, the denoising step, and the prompt, before any reward is observed; it is trained on the reward spread of each rollout group and discarded after training, leaving inference unchanged. On Stable Diffusion 3.5 Medium, EXPLORENET improves held-out GenEval2 by 14% over Flow-GRPO, transfers to two independent compositional benchmarks and five preference and image-quality models, and reaches a 67.2% human preference win-rate. Overall, across our group-relative diffusion RL experiments, we find that exploration is learnable, the shape of the exploration distribution outweighs its magnitude, and rollout quality is more effective than rollout quantity.
Figures & tables
Figure 1: Deterministic flow-matching sampling adds no randomness beyond the initial noise, which limits exploration and leaves the policy ratio without a tractable transition density. Flow-GRPO samples rollouts with isotropic Gaussian perturbations and updates the denoiser fθ with group-relative advantages Ap,i . ExploreNet gϕ predicts the scale map sϕ before each denoising step. The denoiser keeps Lθ , and ExploreNet optimizes Lϕ with the normalized reward spread Bp . Inference uses only the trained denoiser.
Figure 2: With all else fixed, selecting high-spread groups yields the fastest learning.
Figure 3
Figure 4: ExploreNet satisfies the count, color, text, and material constraints that SD3.5-M and Flow-GRPO miss. All three annotators preferred ExploreNet to Flow-GRPO on each prompt.
Figure 5: Held-out GenEval2 during training, avg. over three seeds.
Setting
Method
GenEval2
VLM Judge
GenEval
VVR-Fast
PickScore
HPSv2.1
HPSv3
Aesthetic
ImageReward
Main
Flow-GRPO
0.4982
4.5500
0.6923
0.7624
0.8447
0.3009
8.3142
5.5024
1.1232
Main
Privileged selection
0.5551
4.7500
0.7206
0.7724
0.8463
0.2961
8.1998
5.4916
1.1572
Main
Matched isotropic
0.4372
4.5000
0.7351
0.7791
0.8485
0.3034
8.4229
5.5104
1.1884
Main
ExploreNet
0.5674
4.7000
0.7452
0.8075
0.8509
0.3017
8.6602
5.5997
1.2597
Group size 8
Flow-GRPO
0.4502
4.1000
0.6771
0.7499
0.8478
0.3006
8.2676
5.5168
1.0916
Group size 8
ExploreNet
0.5389
4.3500
0.7422
0.7952
0.8499
0.3100
8.8990
5.6228
1.2806
Table 2: Controls and reduced-budget settings. The main setting uses 24 rollouts per prompt group and 25 denoising steps; the other settings change one of the two. Each row is a single training run; the main-setting Flow-GRPO and ExploreNet rows are the first training seed of Table 3 , so values differ from the three-seed means there.
Figure 7
Figure 8: Channels to which ExploreNet assigns more noise change the image more. Perturbing the two channels with the largest ExploreNet noise scales (middle) changes objects, layout, and background. Perturbing the two with the smallest (right) changes only fine details. Labels give the average noise multiplier ExploreNet assigns to each channel.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Rollouts
Reward calls
Denoising steps
ExploreNet evaluations
Main Flow-GRPO training
24
24
600
0
Main ExploreNet training
24
24
600
600
Privileged selection 48 to 24 training
48
48
1200
0
Matched isotropic training
24
24
600
0
Group size 8 training
8
8
200
200
10 step training
24
24
240
240
Appendix
Table 3: ExploreNet adds gϕ evaluations during training while both trained denoisers retain the same inference cost. Counts describe the sampling passes for one training prompt group or one inference image.
Noise parameterization
Step 1140
Step 3060
Flow-GRPO (isotropic)
0.3825
0.5146
ExploreNet , normalized affine
0.4009
0.5057
ExploreNet , raw clipped
0.4080
0.5358
ExploreNet , normalized and clipped
0.4179
0.5679
Appendix
Table 4: GenEval2 scores for Flow-GRPO and three mappings from the output of gϕ to log standard deviation.
Metric
Smallest
Random
Greatest
GenEval2
0.2579
0.3859
0.4648
VLM Judge
3.8750
4.4625
4.5125
GenEval
0.6171
0.6537
0.6948
VVR-Fast
0.7210
0.7380
0.7614
PickScore
0.8414
0.8433
0.8456
HPSv2.1
0.3016
0.3054
0.3073
Appendix
Table 5: Benchmark results of the three selection rules. Greatest-spread selection is best on eight of nine metrics.
Preferred model
Wins
Ties
Losses
Rate
95% CI
Flow-GRPO over pretrained
118
8
90
56.5%
[48.6,64.4]
ExploreNet over pretrained
143
8
65
68.1%
[60.0,75.9]
ExploreNet over Flow-GRPO
140
7
69
66.4%
[58.6,74.1]
Appendix
Table 6: ExploreNet wins each pairwise preference comparison. Each row contains 216 ratings over 72 prompts. Rates split ties equally, and confidence intervals resample prompts with all associated ratings.
Method
Seed
GenEval2
VLM Judge
GenEval
VVR-Fast
PickScore
HPSv2.1
HPSv3
Aesthetic
ImageReward
Flow-GRPO
1
0.4982
4.5500
0.6923
0.7624
0.8447
0.3009
8.3142
5.5024
1.1232
Flow-GRPO
2
0.4933
4.5750
0.7062
0.7715
0.8467
0.3004
8.3142
5.5184
1.1339
Flow-GRPO
3
0.4956
4.4000
0.6879
0.7598
0.8430
0.2966
8.0051
5.4673
1.0840
ExploreNet
1
0.5674
4.7000
0.7452
0.8075
0.8509
0.3017
8.6602
5.5997
1.2597
ExploreNet
2
0.5972
4.7750
0.7335
0.7928
0.8509
0.3056
8.7869
5.6744
1.2841
ExploreNet
3
0.5337
4.7125
0.7261
0.7703
0.8519
0.3097
8.9658
5.7175
1.2883
Appendix
Table 7: Per-seed results at step 3000 underlying Table 3 . GenEval2 averages three evaluation passes; other metrics use one pass.
Metric
Flow-GRPO
ExploreNet
Gain
GenEval2
0.4809±0.0093
0.5524±0.0300
+0.0715
Appendix
Table 8: The GenEval2 gain persists across the final five checkpoints. Each seed value averages checkpoints 2760 through 3000 over three evaluation passes. Subscripts report sample standard deviation across training seeds.
Evaluation
Paired gain
95% prompt bootstrap interval
GenEval2
+0.0748
[+0.0405,+0.1085]
VVR-Fast
+0.0256
[+0.0202,+0.0311]
Appendix
Table 9: Prompt resampling gives positive gain intervals for GenEval2 and VVRBench-Fast.
Prompt set and metric
Flow-GRPO
ExploreNet
Gain
Pick-a-Pic v1 (500), PickScore
0.8448±0.0019
0.8512±0.0006
+0.0064
HPSv2 anime prompts (800)
0.3130±0.0018
0.3204±0.0038
+0.0073
HPSv2 concept art prompts (800)
0.3034±0.0026
0.3092±0.0054
+0.0058
HPSv2 paintings prompts (800)
0.3038±0.0027
0.3085±0.0039
+0.0047
HPSv2 photo prompts (800)
0.2770±0.0027
0.2845±0.0032
+0.0075
HPSv2.1 macro (3,200)
0.2993±0.0023
0.3056±0.0040
+0.0063
Appendix
Table 10: ExploreNet scores higher on PickScore and on all four HPSv2.1 prompt styles. Bold marks gains whose displayed intervals do not overlap.
Prompt suite
Method
HPSv3
Aesthetic
ImageReward
DrawBench
Flow-GRPO
9.5718±0.2295
5.3664±0.0279
1.0602±0.0503
ExploreNet
10.3331±0.2491
5.5109±0.0625
1.2683±0.0210
PartiPrompts
Flow-GRPO
7.4324±0.1504
5.5783±0.0275
1.2560±0.0110
ExploreNet
7.8012±0.0808
5.7438±0.0507
1.3872±0.0174
DPG-Bench
Flow-GRPO
9.2953±0.2166
5.6607±0.0287
0.8724±0.0316
ExploreNet
9.8481±0.0996
5.8578±0.0401
1.0256±0.0153
Appendix
Table 11: HPSv3, Aesthetic, and ImageReward results on the shared four-suite prompt panel. ExploreNet improves all three metrics on every prompt suite.
Constraint family
Prompts
Flow-GRPO
ExploreNet
Gain
Grounding
302
0.6293±0.0311
0.7683±0.0175
+0.1390
Spatial
665
0.2924±0.0094
0.4187±0.0164
+0.1263
Size
285
0.1816±0.0361
0.2722±0.0270
+0.0906
Topology
242
0.1525±0.0313
0.2385±0.0543
+0.0860
Cardinality
367
0.6790±0.0080
0.6644±0.0568
−0.0146
Appendix
Table 12: ExploreNet improves grounding, spatial, size, and topology constraints on VVRBench-Fast. Each entry is the mean verifier score on the prompts that contain the constraint family, averaged within each training seed and reported as mean and sample standard deviation across the three seeds. Bold marks gains whose displayed intervals do not overlap.
Aggregation
N
Pearson r
Spearman ρ
95% Pearson CI
Prompt and channel combinations
1600
+0.306
+0.332
[+0.273,+0.342]
Centered within prompt
1600
+0.461
+0.490
[+0.438,+0.485]
Channel means
16
+0.598
+0.491
[+0.145,+0.844]
Within prompt, then averaged
100
+0.485
+0.490
[+0.462,+0.506]
Appendix
Table 13: ExploreNet ’s predicted scale tracks intervention sensitivity across prompt and channel aggregation levels. The first two intervals and the last resample prompts with replacement over 20,000 replicates; the channel-means interval is a Fisher interval over the 16 channels. The within-prompt correlation is positive for all 100 prompts, with minimum Spearman ρ=+0.156 .
Denoiser
Bottom quartile
Top quartile
Difference
95% CI
Pretrained
1.521
2.229
+0.708
[+0.271,+0.958]
Flow-GRPO
1.125
1.313
+0.188
[−0.188,+0.521]
ExploreNet
0.771
1.250
+0.479
[+0.250,+0.708]
Pooled
1.139
1.597
+0.458
[+0.319,+0.583]
Appendix
Table 14: Cross denoiser comparison of human change ratings for ExploreNet ’s largest and smallest scale quartiles. The separation is present before training and remains after ExploreNet training. Quartiles are defined within each prompt, and ratings use a scale from 0 to 4.
Perturbed denoiser
Pearson r
Spearman ρ
95% Pearson CI
Within-prompt ρ
Positive prompts
100 prompts of the main analysis
Pretrained
+0.516
+0.488
[+0.027,+0.806]
+0.421
99/100
ExploreNet
+0.598
+0.491
[+0.145,+0.844]
+0.490
100/100
Six prompts of the human study
Pretrained
+0.562
+0.535
[+0.092,+0.827]
+0.451
6/6
Flow-GRPO
+0.509
+0.494
[+0.018,+0.802]
+0.425
6/6
Appendix
Table 15: Correlation between ExploreNet ’s predicted scale and pixel-space change when each denoiser is perturbed. The first three columns use channel means; intervals are Fisher intervals over the 16 channels. Within-prompt ρ averages the per-prompt rank correlation.