Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
Authors: Baoteng Li, Wenzhuo Wu, Kongming Liang, Zhanyu Ma
Organizations: School of Artificial Intelligence, Beijing University of Posts and Telecommunications1 · Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance2
Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.
Figures & tables
Figure 1 : From Visual Jev decisions to generation rewards. GPT-5.4 constructs fixed questions offline from the prompt and ordered references. A frozen, task-trained Visual Jev verifier reads Yes probabilities for each reference-bound condition; their arithmetic mean supplies the GRPO reward. Identity dependence is encoded in each relevant question and its positive-label semantics. Only the generator LoRA is updated during GRPO. The person-and-dog task is illustrative.
Model
Object
Human
HOI
De&Re
Overall
Baseline
41.19
51.05
37.81
40.70
41.78
GRPO (27B)
49.82
60.90
49.71
49.75
51.84
GRPO (Visual Jev)
50.26
61.24
51.46
49.66
52.50
Table 1 : Generation results on the screened MICo-Bench subset. Baseline: official Qwen-Image-Edit-2511. Both GRPO models complete 30 updates. GPT-5.4 evaluates PQ × SC on a 0–100 scale. Overall scores are averages over the same 897 manually selected tasks. One run is reported per method. Higher is better.
Figure 2 : Overall and category-level GRPO results. On the screened 897-task MICo-Bench subset, Visual Jev rewards improve our GPT-5.4 composite score by 10.72 points over the official Qwen-Image-Edit-2511 checkpoint. Relative to direct 27B rewards, the largest category gain is on HOI (+1.75); De&Re is slightly lower ( −0.09 ). Results are observations from one run per method, without a claim of statistical significance.
Figure 3 : Atomic probabilities reveal the basis of a scoring disagreement. All seven questions and four candidates from CP_Object_030 are retained; the displayed English labels abbreviate the reference-bound questions. Direct 27B scores are divided by 100 for display, without implying calibration equivalence to atomic probabilities. The direct scorer prefers C, which receives low Visual Jev probabilities for interface identity and the interface–instrument connection and is ranked last by the human. Visual Jev selects A, matching the human’s first choice, although its full ranking remains imperfect. This post-hoc success case illustrates scoring evidence rather than overall performance.
Method
Acc. ↑
BA ↑
Neg. R ↑
Brier ↓
Untuned Qwen3.5-4B
82.05
84.14
94.67
0.1332
Answer-token SFT
96.05
96.04
96.00
0.0356
BCE
96.15
96.04
95.47
0.0356
BCE + Brier
95.73
95.60
94.93
0.0391
BCE (reward, 136)
–
95.24
–
–
Table 2 : Atomic verification under a fixed training budget. The upper block compares trained methods at update 544 on 936 test examples. The lower row reports the development-selected reward checkpoint at update 136; dashes mark metrics not reported for that checkpoint. Acc.: accuracy; BA: balanced accuracy; Neg. R: negative-class recall. The first three metrics are percentages.
Scoring method
Agreement ↑
Top-1 ↑
Risk ↓
Direct Qwen3.5-4B
59.13
32.54
46.03
Direct Qwen3.5-27B
60.32
39.29
45.24
Visual Jev atomic mean
63.49
42.86
47.62
Table 3 : Offline human preference comparison. All methods use the same 21 valid tasks and 126 strict preference pairs. Direct 4B scoring is an additional control. All values are percentages; identity risk concerns the top-ranked candidate.
Aggregation
Dev. ↑
Val. ↑
Atomic mean (selected)
68.75
63.49
Category mean
68.75
65.08
Category/worst-category mixture
66.67
65.08
Table 4 : Offline aggregation comparison. Development: eight valid tasks and 48 strict preference pairs. Validation: 21 valid tasks and 126 pairs. Values are preference agreement (%). The mixture uses 0.75 times the category mean plus 0.25 times the worst-category mean.