Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts
Organizations: University of Washington
Abstract
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.
Figures & tables
| Dataset | Use | Size | Complexity |
| VVRBench | evaluation | 10,000 | 3–48 |
| VVRBench -Fast | evaluation | 820 | 3–44, 20 each |
| VVRBench -Challenge | evaluation | 720 | 45–80, 20 each |
| VVR-Easy | training | 100,000 | 20 |
| VVR-Matched | training | 100,000 | VVRBench |
| Model | Accuracy (%) | |||||
|---|---|---|---|---|---|---|
| GPT-Image-2 | 86.86 ±0.68 | 97.65 ±0.74 | 98.18 ±0.70 | 91.00 ±1.34 | 81.83 ±1.75 | 65.72 ±2.10 |
| GPT-Image-1-mini | 26.40 ±0.87 | 74.06 ±1.93 | 36.78 ±2.18 | 12.47 ±1.53 | 4.70 ±1.02 | 2.39 ±0.76 |
| FLUX.2-dev | 19.15 ±0.78 | 49.86 ±2.15 | 24.26 ±1.96 | 12.42 ±1.52 | 5.91 ±1.12 | 2.24 ±0.74 |
| HunyuanImage-2.1 | 18.79 ±0.78 | 45.53 ±2.15 | 19.64 ±1.83 | 14.84 ±1.63 | 8.61 ±1.31 | 4.29 ±0.98 |
| Qwen-Image-2512 | 5.79 ±0.47 | 19.16 ±1.75 | 5.87 ±1.14 | 2.11 ±0.73 | 1.10 ±0.56 | 0.15 ±0.29 |
| HiDream-I1-Full | 4.22 ±0.41 | 16.47 ±1.65 | 3.06 ±0.87 | 0.80 ±0.50 | 0.20 ±0.31 | 0.00 ±0.19 |
| Training reward | VVR | GenEval | GenEval2 | OCR | PickScore | HPSv2.1 | HPSv3 | CLIPScore | Aesthetic | ImageReward | UnifiedReward |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Pretrained | 0.028 | 0.616 | 0.237 | 0.476 | 0.841 | 0.300 | 7.689 | 0.956 | 5.517 | 0.929 | 0.636 |
| VVR-Easy | 0.283 | 0.729 | 0.268 | 0.587 | 0.849 | 0.294 | 8.275 | 0.979 | 5.483 | 1.114 | 0.641 |
| +0.255 | +0.113 | +0.031 | +0.111 | +0.008 | +0.586 | +0.023 | +0.185 | +0.005 | |||
| GenEval2 | 0.039 | 0.688 | 0.454 | 0.501 | 0.848 | 0.297 | 8.289 | 0.974 | 5.523 | 1.110 | 0.635 |
| VVR-Easy | 0.218 | 0.718 | 0.478 | 0.532 | 0.848 | 0.302 | 8.511 | 0.976 | 5.533 | 1.155 | 0.637 |
| +0.180 | +0.030 | +0.025 | +0.030 | +0.0004 | +0.005 | +0.222 | +0.002 | +0.009 | +0.045 | +0.0019 |
| VVR win rate vs. baseline | VVR | GenEval2 | GenEval | OCR | DrawBench | Outside VVR |
|---|---|---|---|---|---|---|
| VVR-Easy vs. pretrained | 93.3 | 78.8 | 62.9 | 69.6 | 75.0 | 71.6 |
| [86.7, 98.3] | [67.5, 88.8] | [50.4, 75.0] | [58.3, 80.4] | [65.4, 84.2] | [65.9, 77.1] | |
| GenEval2 VVR-Easy vs. GenEval2 | 89.2 | 62.9 | 57.5 | 56.7 | 57.5 | 58.6 |
| [81.7, 95.4] | [50.8, 74.6] | [45.0, 69.6] | [44.6, 68.8] | [45.4, 69.2] | [52.6, 64.6] |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Exact constraint types & their supported values | Contribution to |
| Grounding | color_attribute : red, orange, yellow, green, cyan, blue, purple, pink | for the referenced group |
| shape_attribute : circle, square, triangle | for the referenced group | |
| color_shape_binding | No additional term; the bound group’s color and shape terms already account for it | |
| Cardinality | exact_count : 1–10 | for the referenced group |
| same_count ; more_than_count ; fewer_than_count | ||
| times_as_many : factor | , |
| Dataset | Size | Candidates | Target distribution |
|---|---|---|---|
| VVRBench | 10,000 | Single strata and compositions of two to six strata at five scene-size settings | Complexity 3–48, capped at each integer complexity |
| VVRBench -Fast | 820 | VVRBench | 20 tasks at each attainable integer complexity from 3 to 44 |
| VVRBench -Challenge | 720 | Scenes seeded by each of the 46 constraint types, with up to six added relation or layout constraints | 20 tasks at each integer complexity from 45 to 80, and at least 20 tasks per non-grounding constraint type |
| VVR-Easy | 100,000 | One constraint type from one stratum | Complexity at most 20, equal quotas over the nine strata |
| VVR-Matched | 100,000 | The VVRBench and VVRBench -Challenge generators | The strata and complexity distribution of VVRBench |
| Open-weight model | Steps | Guidance | API model | Settings |
|---|---|---|---|---|
| FLUX.2-dev | 50 | 4.0 | GPT-Image-2.5-Sunburst (2026-09-08) | medium quality, |
| HunyuanImage-2.1 | 50 | 3.5 | GPT-Image-2 (2026-04-21) | medium quality, |
| Qwen-Image-2512 | 50 | 4.0 | GPT-Image-1-mini | medium quality, |
| HiDream-I1-Full | 50 | 5.0 | Gemini-3-Pro-Image | 1K |
| FLUX.1-dev | 28 | 3.5 | Gemini-3.1-Flash-Image | 1K |
| FLUX.1-schnell | 4 | 0.0 | Gemini-3.1-Flash-Lite-Image | 1K |
| Model | Accuracy (%) | 3 to 10 | 11 to 18 | 19 to 26 | 27 to 35 | 36 to 44 |
|---|---|---|---|---|---|---|
| GPT-Image-2.5-Sunburst | 84.51 ±2.64 | 100.00 ±2.67 | 98.75 ±3.19 | 96.25 ±4.19 | 77.22 ±6.66 | 56.67 ±7.30 |
| GPT-Image-2 | 82.20 ±2.77 | 98.57 ±3.63 | 98.75 ±3.19 | 95.62 ±4.38 | 74.44 ±6.84 | 50.56 ±7.24 |
| Gemini-3.1-Flash-Image | 47.20 ±3.42 | 85.71 ±6.75 | 60.00 ±7.74 | 51.25 ±7.68 | 30.56 ±7.08 | 18.89 ±6.35 |
| Gemini-2.5-Flash-Image | 45.61 ±3.42 | 76.43 ±7.68 | 65.62 ±7.65 | 51.25 ±7.68 | 31.11 ±7.10 | 13.33 ±5.74 |
| Gemini-3.1-Flash-Lite-Image | 42.07 ±3.41 | 85.71 ±6.75 | 56.25 ±7.74 | 40.00 ±7.74 | 25.00 ±6.80 | 14.44 ±5.88 |
| Gemini-3-Pro-Image | 37.20 ±3.36 | 62.14 ±8.26 | 52.50 ±7.71 | 38.75 ±7.73 | 26.67 ±6.90 | 13.33 ±5.74 |
| Model | Overall | Grounding | Cardinality | Spatial | Size | Topology |
|---|---|---|---|---|---|---|
| GPT-Image-2 | 86.9 | 86.8 | 80.0 | 85.4 | 87.6 | 85.6 |
| GPT-Image-1-mini | 26.4 | 26.6 | 14.2 | 20.7 | 19.2 | 10.7 |
| FLUX.2-dev | 19.1 | 19.0 | 8.5 | 16.3 | 11.7 | 9.1 |
| HunyuanImage-2.1 | 18.8 | 19.0 | 10.3 | 16.5 | 6.5 | 11.5 |
| Qwen-Image-2512 | 5.8 | 6.7 | 2.2 | 4.4 | 2.4 | 3.3 |
| HiDream-I1-Full | 4.2 | 5.3 | 2.0 | 2.6 | 1.6 | 1.5 |
| Setting | Value |
|---|---|
| Trainable parameters | LoRA on the eight attention projections add_k , add_q , add_v , add_out , k , q , v , and out ; rank 32 and . |
| Generation | pixels; 25 denoising steps; classifier-free guidance 4.5; Gaussian sampling noise with level 0.7. |
| Rollout batch | 32 prompt groups per update, 24 rollouts per prompt, and 768 generated images per update. |
| Flow GRPO | One inner epoch; advantages centered within each prompt group and divided by the standard deviation over the complete rollout batch; advantages clipped to ; policy-ratio clip ; KL coefficient 0.04. The first 24 of the 25 sampled transitions contribute to the update. |
| Optimization | AdamW; learning rate ; , ; ; weight decay ; maximum gradient norm 1.0; FP16 mixed precision with TF32 enabled; exponential moving average. |
| Training duration | 3,000 optimizer updates, corresponding to 96,000 prompt groups and 2,304,000 generated images. |
| Training reward | Accuracy (%) | |||||
|---|---|---|---|---|---|---|
| SD3.5-M (pretrained) | 2.81 ±0.34 | 12.01 ±1.47 | 1.19 ±0.59 | 0.35 ±0.37 | 0.05 ±0.23 | 0.00 ±0.19 |
| GenEval2 | 3.87 ±0.40 | 15.42 ±1.61 | 2.39 ±0.78 | 0.91 ±0.52 | 0.05 ±0.23 | 0.05 ±0.23 |
| GenEval2 + VVR-Easy | 21.82 ±0.82 | 54.51 ±2.15 | 32.05 ±2.12 | 14.08 ±1.60 | 6.41 ±1.16 | 1.10 ±0.56 |
| VVR-Easy | 28.27 ±0.89 | 67.72 ±2.04 | 45.45 ±2.23 | 17.51 ±1.73 | 8.36 ±1.29 | 1.35 ±0.60 |
| VVR-Matched | 46.60 ±0.98 | 67.68 ±2.04 | 59.17 ±2.21 | 45.62 ±2.20 | 38.39 ±2.15 | 21.82 ±1.86 |
| GenEval2 + VVR-Matched | 33.50 ±0.93 | 57.93 ±2.13 | 47.27 ±2.23 | 32.34 ±2.09 | 20.32 ±1.82 | 9.22 ±1.34 |
| Counts | Relations | |||||
|---|---|---|---|---|---|---|
| Range | Model | Partial | All | Partial | All | Accuracy (%) |
| Pretrained | 0.76 | 0.46 | 0.29 | 0.19 | 12.01 | |
| VVR-Easy | 0.98 | 0.92 | 0.80 | 0.58 | 67.72 | |
| VVR-Matched | 0.98 | 0.93 | 0.80 | 0.57 | 67.68 | |
| Pretrained | 0.70 | 0.17 | 0.25 | 0.13 | 1.19 | |
| VVR-Easy | 0.96 | 0.79 | 0.75 | 0.51 | 45.45 | |
| Training reward | VVR | GenEval | GenEval2 | OCR | PickScore | HPSv2.1 | HPSv3 | CLIPScore | Aesthetic | ImageReward | UnifiedReward |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Pretrained | 0.028 | 0.616 | 0.237 | 0.476 | 0.841 | 0.300 | 7.689 | 0.956 | 5.517 | 0.929 | 0.636 |
| GenEval2 | 0.039 | 0.688 | 0.454 | 0.501 | 0.848 | 0.297 | 8.289 | 0.974 | 5.523 | 1.110 | 0.635 |
| GenEval2 + VVR-Easy | 0.218 | 0.718 | 0.478 | 0.532 | 0.848 | 0.302 | 8.511 | 0.976 | 5.533 | 1.155 | 0.637 |
| VVR-Easy | 0.283 | 0.729 | 0.268 | 0.587 | 0.849 | 0.294 | 8.275 | 0.979 | 5.483 | 1.114 | 0.641 |
| add VVR-Easy | +0.180 | +0.030 | +0.025 | +0.030 | +0.0004 | +0.005 | +0.222 | +0.002 | +0.009 | +0.045 | +0.002 |
| VVR-Matched | 0.466 | 0.709 | 0.360 | 0.536 | 0.847 | 0.291 | 7.997 | 0.980 | 5.472 | 1.109 | 0.636 |
| Contrast | PickScore | HPSv2.1 | CLIPScore | Aesthetic | ImageReward |
|---|---|---|---|---|---|
| GenEval2 VVR-Easy GenEval2 | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| GenEval2 VVR-Matched GenEval2 | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| OCR VVR-Easy OCR | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| Five-reward VVR-Easy Five-reward | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| Annotator pair | Agreement, both chose (%) | Agreement, ties as a label (%) | Cohen’s |
|---|---|---|---|
| 1 vs. 2 | 84.9 (298/351) | 76.0 (304/400) | 0.50 |
| 1 vs. 3 | 81.9 (276/337) | 72.0 (288/400) | 0.44 |
| 2 vs. 3 | 84.6 (312/369) | 78.8 (315/400) | 0.53 |
| Mean | 83.8 | 75.6 | 0.49 |