Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.
Figures & tables
Figure 1 : Thinking Reward Model (TRM) follows the “Think Before You Score” paradigm. Given a visual generation case, TRM first formulates case-adaptive rubrics that specify what matters , then performs rubric-guided assessment before producing the final pointwise reward. This unified process achieves strong reward-modeling performance across image generation and editing, and effectively guides reinforcement learning for diverse visual generation models.
Method
Task
Modeling Paradigm
Scoring
Adaptive Rubrics
Fine-Grained Verification
RL Optimization
Point
Pair
ImageReward
T2I
Regressive
✓
–
✗
✗
✗
UnifiedReward
T2I, T2V
Generative
✓
✓
✗
✓
✗
RationalRewards
TI2I, T2I
Generative
✓
✓
✗
✓
✗
RewardDance
T2I
Generative
–
✓
✗
✓
✗
FIRM-Reward
TI2I, T2I
Generative
✓
–
✗
✓
✗
Table 1 : Comparison of representative reward models.
Figure 2 : Overview of data construction and training pipeline. We construct diverse and balanced data for image generation and editing, annotate rubrics and scores through a two-stage process, learn rubric-then-score evaluation via SFT, and further improve pointwise scoring with PD-GRPO using difficulty-aware pairwise preference supervision.
Model
Size
GenAI-T2I
MMRB2-T2I
Proprietary Models
GPT-4.1
–
60.5
65.8
Gemini 2.5 Flash
–
65.8
63.1
Gemini 2.5 Pro
–
66.2
70.5
Gemini 3 Pro
–
73.1
74.4
Open-Source Models
Table 2 : Performance on image generation reward-modeling benchmarks.
Model
Size
EditScore-ERB
MMRB2
EditReward-ERB
EditReward-Compass
IF
VC
O
2-path
IA
VC
Proprietary Models
GPT-4.1
–
0.673
0.602
0.705
68.2
72.1
0.747
0.485
GPT-5
–
0.777
0.669
0.755
73.8
73.0
–
–
Gemini 2.5 Pro
–
0.703
0.560
0.722
71.3
78.3
–
–
Gemini 3.1 Pro
–
0.877
0.716
0.841
74.9
73.9
0.832
0.600
Table 3 : Performance on image editing reward-modeling benchmarks.
Model
GenEval ↑
DPG-Bench ↑
TIIF-Short ↑
TIIF-Long ↑
Representative Image Generation Models
OmniGen2
0.80
83.60
70.20
70.30
LongCat-Image
0.87
86.80
–
–
Qwen-Image
0.87
88.32
86.14
86.83
LLaDA-Image
0.85
87.48
–
–
Z-Image
0.84
88.14
80.20
83.01
Table 4 : Results of TRM-guided RL on image generation.
Model
ImgEdit
GEdit-Bench-EN
GEdit-Bench-CN
G_SC
G_PQ
G_O ↑
G_SC
G_PQ
G_O ↑
Representative Image Editing Models
OmniGen2
3.44
7.16
6.77
6.41
–
–
–
LongCat-Image-Edit
4.44
8.13
8.18
7.75
8.14
8.12
7.73
Qwen-Image-Edit2509
4.34
7.97
7.71
7.48
7.99
7.68
7.47
LLaDA-Image
–
8.04
7.18
7.34
7.71
7.59
7.29
Table 5 : Results of TRM-guided RL on image editing.
Figure 3 : Qualitative comparison of SenseNova-U1.5 before and after RL fine-tuning with TRM.
Image Generation: BAGEL
Reward / Method
GenEval ↑
DPG-Bench ↑
TIIF-S ↑
TIIF-L ↑
Base
0.86
85.07
74.91
75.62
AlphaGRPO
0.86
85.10
77.70
78.10
TRM (Ours)
0.89
86.60
80.68
81.43
Image Editing: SenseNova-U1.5
Reward Model
ImgEdit
GEdit-EN
GEdit-CN
RM Size
Table 6 : Ablation and controlled comparisons of reward-guided optimization.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Categories
Evaluation Focus
Alignment
subject, attribute, count space, action scene
Subject presence, attribute binding, counting, spatial relations, actions, and scene conditions
Text
few, many, combo
Short text, long text, and text rendering combined with other visual requirements
Reasoning
logic, causal, analog
Logical, causal, behavioral, analogical, generalization, and procedural reasoning
Aesthetics
quality, color, view, detail
Overall quality, color, composition/viewpoint, and fine-grained detail
Others
portrait, body anomaly, poster, multiling text
Portrait realism, body/hand anomalies, poster composition, and multilingual text rendering
Appendix
Table 7 : Taxonomy of the image generation SFT data. We organize diverse generation cases into five complementary groups for data balancing; these categories do not serve as fixed evaluation rubrics.
Category
Evaluation Focus
Addition
Accurate insertion and natural integration
Remove
Complete removal and seamless region restoration
Replace
Accurate replacement and natural integration
Text Editing
Text accuracy, legibility, and visual consistency
Background Change
Background accuracy and foreground preservation
Style Transfer
Target style alignment and content preservation
Appendix
Table 8 : Taxonomy of the image editing SFT data. The collected cases cover both direct visual modifications and reasoning-intensive editing operations.
Figure 4 : Task distributions of the SFT data. Left: the image-generation SFT data are organized into fine-grained task categories under the compact taxonomy in Table 7 . Right: the image-editing SFT data span 19 fine-grained task categories under the taxonomy in Table 8 . Together, these distributions illustrate the balanced and diverse coverage of our SFT data across visual generation tasks.
Model
Avg. Acc.
Forward Acc.
Reverse Acc.
Consistent
Inconsistent
Qwen3-VL-8B
62.0
63.9
59.9
55.9
44.1
Qwen3.5-9B
51.4
51.3
51.5
45.3
54.7
Qwen2.5-VL-72B
65.8
65.8
65.8
74.6
25.5
Appendix
Table 9 : Pairwise preference consistency under order reversal on MMRB2. Forward and reverse evaluations present the same candidates in opposite orders. Consistency is measured after mapping predictions back to candidate identities. All values are percentages.
Figure 5 : Reward dynamics during TRM-guided RL. Training reward, reward standard deviation, and evaluation reward for SenseNova-U1.5 and BAGEL on image editing and BAGEL on image generation.
Figure 6 : Training dynamics with different reward models. Comparison of mean reward and reward standard deviation during RL guided by TRM and the larger EditScore-72B reward model.
Figure 7 : Comparison of score distributions between PD-GRPO and Bradley–Terry optimization.
Figure 8 : Qualitative results of TRM-guided optimization on BAGEL for image editing. For each example, we show the input image, the output of the original model,and the output after reinforcement learning with TRM as the reward. TRM-guided optimization improves instruction following and edit quality across diverse editing tasks while preserving unrelated image content.
Model
GenAI-T2I
MMRB2-T2I
Qwen3.5-9B (Baseline)
54.6
53.5
TRM(SFT)
67.7
62.8
TRM(RL)
68.4
63.9
Appendix
Table 10 : Tie-aware evaluation on image-generation reward-modeling benchmarks. All successfully parsed candidate pairs are retained, and a predicted tie is assigned 0.5 credit. We report pairwise preference accuracy (%).
Figure 9 : Qualitative results of TRM-guided optimization on SenseNova-U1.5 for image editing. We compare outputs from the original model and its TRM-optimized counterpart across diverse editing instructions. TRM-guided optimization produces more accurate edits while maintaining visual consistency with the input image.
Figure 10 : Qualitative results of TRM-guided optimization on BAGEL for image generation. Given the same text prompts, we compare generations from the original model and its TRM-optimized counterpart. TRM-guided optimization improves prompt alignment and fine-grained visual fidelity across diverse generation cases.
Figure 11 : Qualitative results of TRM-guided optimization on FLUX.1-dev for image generation. We compare generations from the original model and the model optimized using TRM as the reward. The optimized model better satisfies prompt requirements while preserving overall visual quality.
Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric scores than as a deterministic scalar. Existing scalar, score-token, and pairwise reward models over-compress uncertainty and fine-grained score differences, while reasoning-based generative rewards provide stronger judgments but are costly to deploy and difficult to use as direct optimization signals. We propose Z-Reward, a teacher-student reward modeling framework that decouples reasoning-heavy judgment from efficient reward deployment. The teacher is a large VLM that uses reasoning to infer rubric-aligned score distributions, and is trained with Group-wise Direct Score Optimization (GDSO), which combines policy-gradient rewards from distribution expectations with direct pointwise and pairwise supervision on score distributions and score gaps. The student is trained with Reasoning-Internalized Score Distillation (RISD), which transfers the teacher's reasoning-conditioned score distribution into a compact VLM without requiring explicit reasoning chains at inference time. On our internally annotated evaluation set, the 27B GDSO teacher reaches 89.6% human preference accuracy, outperforming SFT, RewardDance, and GRPO, while the 9B RISD student reaches 88.6%, outperforming the OPD baseline and closely matching the larger teacher. We further show that Z-Reward can serve as a differentiable reward signal for text-to-image optimization, yielding a 41.3% net human-preference improvement over the SFT baseline.
Xin Jin, Huanqia Cai, Zhen Li +9
Z-Image Team, Alibaba Group · VCIP, CS, Nankai University
Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies. To address these limitations, we present a novel framework that finetunes generative models using distribution-wise rewards, ensuring better alignment with real-world data distributions. Unlike rewards that evaluate samples individually, distribution-wise reward accounts for the data distribution of the samples, mitigating the mode collapse problem that occurs when all samples optimize towards the same direction independently. To overcome the prohibitive computational cost of estimating these rewards, we introduce a subset-replace strategy that efficiently provides reward signals by updating only a small subset of a generated reference set. Additionally, we apply RL to optimize post-hoc model merging coefficients, potentially mitigating the train-inference inconsistency caused by introducing stochastic differential equation (SDE) in regular RL practices. Extensive experiments show our approach significantly improves FID-50K across various base models, from 8.30 to 5.77 for SiT and from 3.74 to 3.52 for EDM2. Qualitative evaluation also confirms that our method enhances perceptual quality while preserving sample diversity.
Ruihang Li, Mengde Xu, Shuyang Gu +4
University of Science and Technology of China · Shanghai Innovation Institute · Hunyuan Frontier Lab, Tencent +1
Aligning large visual generative models with human feedback is often performed through pairwise preference optimization. While such approaches are conceptually simple, they fundamentally rely on annotated pairs, limiting scalability in settings where feedback is collected as independent scalar ratings. In this work, we revisit the KL-regularized alignment objective and show that the optimal policy implicitly compares each sample's reward to an instance-specific baseline that is generally intractable. We propose a threshold-guided alignment framework that replaces this oracle baseline with a data-driven global threshold estimated from empirical score statistics. This formulation turns alignment into a binary decision task on unpaired data, enabling effective optimization directly from scalar feedback. We also incorporate a confidence weighting term to emphasize samples whose scores deviate strongly from the threshold, improving sample efficiency. Experiments across both diffusion and masked generative paradigms, spanning three test sets and five reward models, show that our method consistently improves preference alignment over previous methods. These results position our threshold-guided framework as a simple yet principled alternative for aligning visual generative models without paired comparisons.
Jinbin Bai, Yu Lei, Qingyu Shi +6
National University of Singapore · Collov Labs · Peking University +2