Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model's training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator's current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.
Figures & tables
Figure 1: Latent reward hacking. We fine-tune Wan2.1-T2V-1.3B against a fixed latent reward model. Top: frames generated, labeled with the image quality. As the reward rises from 0.40 to 0.86 and stays well above its initial value, the video degrades and loses detail. Bottom: mean latent reward and imaging quality over training. The reward rises while imaging quality declines.
Figure 2: Latent reward hacking and co-evolving supervision. With a fixed reward model, generator optimization can move samples beyond the region covered by reward supervision and exploit unsupported high scores (left). With co-evolution, reward supervision follows the generator distribution as it changes (right).
Step
Reward
Dyn. deg.
Imaging
Subj. cons.
sreal
sgen
Spread
0
0.636
68.06
67.91
96.35
0.551
0.666
1.16
500
0.673
54.17
67.07
97.04
0.554
0.753
0.76
1000
0.742
27.78
64.23
97.59
0.540
0.781
0.74
Table 1: Fixed latent reward optimization. Reward is the frozen LRM mean sigmoid score. sreal and sgen are mean scores of 256 real and 256 generated latents, respectively, evaluated at t=591 for each checkpoint. Spread is the ratio of within-generated to within-real 10-nearest-neighbor distances in normalized reward features; lower values indicate more concentrated. Blue marks increasing generated scores; Red marks declines in corresponding degree.
Figure 3: Overall framework of CoRe .
Figure 4: Generated and real scores under fixed and co-evolving rewards. (a) Gap between the mean scores of generated and real latents over training; the horizontal axis shows normalized progress for each run. Under the fixed reward, the gap grows, while under CoRe it stays near zero. (b, c) Histograms of reward-model scores for real and generated latents at the end of training.
Figure 5: Reward-model features of real and generated latents at t=591 , visualized with t-SNE and PCA. Each panel shows: coloured by domain (left; blue = real, red = generated) and by the reward model’s score (right; brighter = higher). Under the fixed reward, generated latents form a separate cluster that receives uniformly high scores (mean score 0.82 for generated vs. 0.55 for real); under ours, generated latents overlap with real data and their scores no longer depend on domain.
Figure 6: Qualitative comparison .
Table 8
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Real vs. generated samples in the reward-feature space under a fixed reward model , colored by reward score ( ▲ real, ∙ generated). The two domains are almost perfectly separated, and generated samples concentrate in a high-reward region (yellow) beyond the range covered by real samples.
Figure 8: Training reward against a fixed latent reward model. (a) Without regularization, the reward saturates near 0.95 within about 20 updates, while the decoded videos blur by step 15, show grid-like artifacts by step 24, and collapse by step 100. (b) With the flow-matching anchor, the reward no longer saturates but also shows no clear upward trend over 2.5k updates. A fixed reward thus leaves two outcomes: rapid exploitation, or little effective optimization.
DiT block
Preference acc. ↑
Reward–quality corr. ↑
Stable steps ↑
4
68.2
0.41
420
8
76.4
0.54
850
12
74.1
0.43
510
16
71.6
0.37
290
Appendix
Table 6: Reward models built on different DiT blocks. Preference acc.: held-out accuracy on VideoDPO pairs. Reward–quality corr.: Pearson correlation between reward scores and VBench imaging quality on generated videos. Stable steps: number of generator updates before dynamic degree falls below 0.3 .
Experiment
Ckpt
SC
BC
MS
DD
AQ
IQ
Fixed LRM
500
0.969
0.972
0.983
0.681
0.642
0.600
v2
300
0.956
0.961
0.969
0.896
0.596
0.638
v2
400
0.966
0.967
0.975
0.778
0.619
0.656
step200
200
0.964
0.970
0.987
0.681
0.603
0.679
v2
750
0.978
0.982
0.992
0.194
0.631
0.629
w1p5
350
0.946
0.960
0.984
0.837
0.567
0.627
Appendix
Table 7: VBench scores for all runs. SC = subject_consistency, BC = background_consistency, MS = motion_smoothness, DD = dynamic_degree, AQ = aesthetic_quality, IQ = imaging_quality.
Objective
Step
Dyn.
Imaging
Aesth.
Subj.
Smooth
Absolute target
450
61.11
66.13
61.05
97.14
98.19
500
56.94
64.09
61.24
96.74
98.28
Relative gap
450
72.22
64.96
60.47
96.25
98.14
500
69.44
65.85
59.23
95.88
97.69
Appendix
Table 8: Generator objective, evaluated at two checkpoints of the same 100-step run.
Achieving simultaneous preference alignment and distillation acceleration in video diffusion models remains an open challenge. Existing methods optimize the two objectives over mismatched representation spaces, where improving one objective often compromises the other. To overcome this, we propose Reward Lightning, a unified framework that aligns and accelerates a video diffusion model within a single shared representation. Its central principle is homology: both objectives are evaluated on identical latent features, which mitigates the gradient conflicts that arise when they are optimized over disjoint representations. As a foundational component, we first introduce a latent reward model (LRM) that scores videos directly in the latent space, without decoding back to the pixel space. Building on the LRM, homologous preference distillation (HPD) reuses this shared backbone to perform adversarial distillation and preference alignment jointly, yielding few-step generators that remain faithful and well aligned. Extensive experiments demonstrate that the LRM surpasses pixel-level and latent-level reward baselines by 11.0% and 14.7% in preference accuracy, and that Reward Lightning generates high-fidelity videos in merely 1 to 4 steps, improving the average VBench score by 2.1% while leading in text alignment, motion quality, and visual quality. Project page: https://reward-lightning.github.io.
Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. However, existing latent reward models output only scalar scores. They do not estimate the uncertainty of each prediction. The generator therefore cannot determine which feedback is reliable. This can drive optimization in the wrong direction and lead to reward hacking. We propose \textsc{SURE}, a unified latent-space framework for image and video diffusion models. It learns reward distributions and directly uses their reliability to guide dense post-training. First, we propose sample-adaptive latent reward model (\textsc{SURE-LRM}). It predicts a Gaussian utility for each noisy latent. Its mean predicts the reward score. Its variance reflect the uncertainty of prediction without human annotation. The learned distribution then guides post-training through uncertainty-guided reward feedback learning (\textsc{SURE-REFL}). This method provides uncertainty-guided dense feedback along the denoising trajectory. At selected transitions, \textsc{SURE-REFL} queries the frozen \textsc{SURE-LRM}. It converts detached variance into reliability weights for samples at the same transition. Each weighted reward is backpropagated only through its local transition. The entire process remains in latent space and requires neither pixel-space decoding nor the full denoising graph. Experiments show that \textsc{SURE-LRM} improves preference prediction over strong baselines. \textsc{SURE-REFL} achieves the sota performance among various metrics and further improves optimization stability. It also achieves the highest VBench quality, semantic, and total scores among the evaluated methods.
Rui Li, Yuanzhi Liang, Ke Hao +4
University of Science and Technology of China, Hefei, China · Institute of Artificial Intelligence, China Telecom (TeleAI), Shanghai, China · Shanghai Jiao Tong University, Shanghai, China
Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward models, pre-trained for semantic alignment, struggle to capture the essential perceptual qualities-such as aesthetics, composition, and visual harmony. In this work, we argue that a model capable of high-fidelity generation must possess a profound understanding of these visual attributes. Based on this insight, we introduce the Diffusion-based Reward Model (DRM), a novel paradigm that use the pre-trained diffusion model as a powerful evaluative backbone. A key advantage of the DRM is its unique ability to assess not only the final image but also the noisy intermediate latents at any stage of the generative process. We leverage this step-wise evaluative capacity in two ways. First, we propose Step-wise GRPO, a reinforcement learning algorithm that provides dense, per-step rewards to resolve the imprecise credit assignment problem in GRPO algorithm, leading to more stable and effective alignment. Second, we introduce Step-wise Sampling, a novel inference strategy that employs the DRM as a dynamic guide to evaluate multiple generation paths at each step, steering the process towards higher-quality outcomes. Extensive experiments confirm that our approach significantly enhances the final quality of generated images. Code: https://github.com/jjaxonx/DRM.
Jaxon Zhang, Binxin Yang, Hubery Yin +2
Peking University · Work done during an internship at WeChat Vision, Tencent Inc. · WeChat Vision, Tencent Inc.