Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
Authors: Baoteng Li, Wenzhuo Wu, Kongming Liang, Zhanyu Ma
Organizations: School of Artificial Intelligence, Beijing University of Posts and Telecommunications1 · Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance2
Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.
Figures & tables
Figure 1 : From Visual Jev decisions to generation rewards. GPT-5.4 constructs fixed questions offline from the prompt and ordered references. A frozen, task-trained Visual Jev verifier reads Yes probabilities for each reference-bound condition; their arithmetic mean supplies the GRPO reward. Identity dependence is encoded in each relevant question and its positive-label semantics. Only the generator LoRA is updated during GRPO. The person-and-dog task is illustrative.
Model
Object
Human
HOI
De&Re
Overall
Baseline
41.19
51.05
37.81
40.70
41.78
GRPO (27B)
49.82
60.90
49.71
49.75
51.84
GRPO (Visual Jev)
50.26
61.24
51.46
49.66
52.50
Table 1 : Generation results on the screened MICo-Bench subset. Baseline: official Qwen-Image-Edit-2511. Both GRPO models complete 30 updates. GPT-5.4 evaluates PQ × SC on a 0–100 scale. Overall scores are averages over the same 897 manually selected tasks. One run is reported per method. Higher is better.
Figure 2 : Overall and category-level GRPO results. On the screened 897-task MICo-Bench subset, Visual Jev rewards improve our GPT-5.4 composite score by 10.72 points over the official Qwen-Image-Edit-2511 checkpoint. Relative to direct 27B rewards, the largest category gain is on HOI (+1.75); De&Re is slightly lower ( −0.09 ). Results are observations from one run per method, without a claim of statistical significance.
Figure 3 : Atomic probabilities reveal the basis of a scoring disagreement. All seven questions and four candidates from CP_Object_030 are retained; the displayed English labels abbreviate the reference-bound questions. Direct 27B scores are divided by 100 for display, without implying calibration equivalence to atomic probabilities. The direct scorer prefers C, which receives low Visual Jev probabilities for interface identity and the interface–instrument connection and is ranked last by the human. Visual Jev selects A, matching the human’s first choice, although its full ranking remains imperfect. This post-hoc success case illustrates scoring evidence rather than overall performance.
Method
Acc. ↑
BA ↑
Neg. R ↑
Brier ↓
Untuned Qwen3.5-4B
82.05
84.14
94.67
0.1332
Answer-token SFT
96.05
96.04
96.00
0.0356
BCE
96.15
96.04
95.47
0.0356
BCE + Brier
95.73
95.60
94.93
0.0391
BCE (reward, 136)
–
95.24
–
–
Table 2 : Atomic verification under a fixed training budget. The upper block compares trained methods at update 544 on 936 test examples. The lower row reports the development-selected reward checkpoint at update 136; dashes mark metrics not reported for that checkpoint. Acc.: accuracy; BA: balanced accuracy; Neg. R: negative-class recall. The first three metrics are percentages.
Scoring method
Agreement ↑
Top-1 ↑
Risk ↓
Direct Qwen3.5-4B
59.13
32.54
46.03
Direct Qwen3.5-27B
60.32
39.29
45.24
Visual Jev atomic mean
63.49
42.86
47.62
Table 3 : Offline human preference comparison. All methods use the same 21 valid tasks and 126 strict preference pairs. Direct 4B scoring is an additional control. All values are percentages; identity risk concerns the top-ranked candidate.
Aggregation
Dev. ↑
Val. ↑
Atomic mean (selected)
68.75
63.49
Category mean
68.75
65.08
Category/worst-category mixture
66.67
65.08
Table 4 : Offline aggregation comparison. Development: eight valid tasks and 48 strict preference pairs. Validation: 21 valid tasks and 126 pairs. Values are preference agreement (%). The mixture uses 0.75 times the category mean plus 0.25 times the worst-category mean.
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.
Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.
Multi-subject personalized image generation requires the precise rendering of all requested reference identities and their specified interactions based on a guiding prompt. However, state-of-the-art models still struggle with this process, frequently omitting subjects, failing to preserve reference appearances, or misattributing interactions. Furthermore, existing metrics designed primarily for single-subject fidelity cannot reliably capture these errors, suffering severe degradation in ranking separability and failing to align with human preference as the subject count increases. To address this gap, we introduce Multi-subject Interaction Benchmark and Evaluator (MIBE), a unified framework comprising a Multi-subject Interaction Benchmark (MIB) and a Multi-subject Interaction Evaluator (MIE). MIB systematically covers diverse relation types and scene complexities through a decoupled data regime. This consists of a 60K-pair VLM-labeled Silver Set for scalable metric training and a 4K-pair double-blind Human Evaluation Gold Set covering a diverse range of state-of-the-art generators, with the Silver Set reaching 95.1% cross-VLM preference agreement. To demonstrate the utility of this benchmark, we present MIE, a lightweight, reference-conditioned evaluator trained exclusively on the Silver Set with a dual-head ranking and diagnosis objective. MIE exhibits strong cross-generator generalization on the Gold Set, achieving 0.922 overall pairwise accuracy against human preference, including 0.982 on seen generators and 0.884 on unseen generators. By outperforming a broad spectrum of baseline metrics, including CLIP and DINO variants, MIE demonstrates that diagnostic supervision can preserve ranking separability and human alignment where traditional evaluators collapse.
Zhihan Chen, Yuhuan Zhao, Yijie Zhu +7
University of California, Los Angeles · University of Southern California · DeerLab LLC +5