Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model's training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator's current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.
Figures & tables
Figure 1: Latent reward hacking. We fine-tune Wan2.1-T2V-1.3B against a fixed latent reward model. Top: frames generated, labeled with the image quality. As the reward rises from 0.40 to 0.86 and stays well above its initial value, the video degrades and loses detail. Bottom: mean latent reward and imaging quality over training. The reward rises while imaging quality declines.
Figure 2: Latent reward hacking and co-evolving supervision. With a fixed reward model, generator optimization can move samples beyond the region covered by reward supervision and exploit unsupported high scores (left). With co-evolution, reward supervision follows the generator distribution as it changes (right).
Step
Reward
Dyn. deg.
Imaging
Subj. cons.
sreal
sgen
Spread
0
0.636
68.06
67.91
96.35
0.551
0.666
1.16
500
0.673
54.17
67.07
97.04
0.554
0.753
0.76
1000
0.742
27.78
64.23
97.59
0.540
0.781
0.74
Table 1: Fixed latent reward optimization. Reward is the frozen LRM mean sigmoid score. sreal and sgen are mean scores of 256 real and 256 generated latents, respectively, evaluated at t=591 for each checkpoint. Spread is the ratio of within-generated to within-real 10-nearest-neighbor distances in normalized reward features; lower values indicate more concentrated. Blue marks increasing generated scores; Red marks declines in corresponding degree.
Figure 3: Overall framework of CoRe .
Figure 4: Generated and real scores under fixed and co-evolving rewards. (a) Gap between the mean scores of generated and real latents over training; the horizontal axis shows normalized progress for each run. Under the fixed reward, the gap grows, while under CoRe it stays near zero. (b, c) Histograms of reward-model scores for real and generated latents at the end of training.
Figure 5: Reward-model features of real and generated latents at t=591 , visualized with t-SNE and PCA. Each panel shows: coloured by domain (left; blue = real, red = generated) and by the reward model’s score (right; brighter = higher). Under the fixed reward, generated latents form a separate cluster that receives uniformly high scores (mean score 0.82 for generated vs. 0.55 for real); under ours, generated latents overlap with real data and their scores no longer depend on domain.
Figure 6: Qualitative comparison .
Table 8
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Real vs. generated samples in the reward-feature space under a fixed reward model , colored by reward score ( ▲ real, ∙ generated). The two domains are almost perfectly separated, and generated samples concentrate in a high-reward region (yellow) beyond the range covered by real samples.
Figure 8: Training reward against a fixed latent reward model. (a) Without regularization, the reward saturates near 0.95 within about 20 updates, while the decoded videos blur by step 15, show grid-like artifacts by step 24, and collapse by step 100. (b) With the flow-matching anchor, the reward no longer saturates but also shows no clear upward trend over 2.5k updates. A fixed reward thus leaves two outcomes: rapid exploitation, or little effective optimization.
DiT block
Preference acc. ↑
Reward–quality corr. ↑
Stable steps ↑
4
68.2
0.41
420
8
76.4
0.54
850
12
74.1
0.43
510
16
71.6
0.37
290
Appendix
Table 6: Reward models built on different DiT blocks. Preference acc.: held-out accuracy on VideoDPO pairs. Reward–quality corr.: Pearson correlation between reward scores and VBench imaging quality on generated videos. Stable steps: number of generator updates before dynamic degree falls below 0.3 .
Experiment
Ckpt
SC
BC
MS
DD
AQ
IQ
Fixed LRM
500
0.969
0.972
0.983
0.681
0.642
0.600
v2
300
0.956
0.961
0.969
0.896
0.596
0.638
v2
400
0.966
0.967
0.975
0.778
0.619
0.656
step200
200
0.964
0.970
0.987
0.681
0.603
0.679
v2
750
0.978
0.982
0.992
0.194
0.631
0.629
w1p5
350
0.946
0.960
0.984
0.837
0.567
0.627
Appendix
Table 7: VBench scores for all runs. SC = subject_consistency, BC = background_consistency, MS = motion_smoothness, DD = dynamic_degree, AQ = aesthetic_quality, IQ = imaging_quality.
Objective
Step
Dyn.
Imaging
Aesth.
Subj.
Smooth
Absolute target
450
61.11
66.13
61.05
97.14
98.19
500
56.94
64.09
61.24
96.74
98.28
Relative gap
450
72.22
64.96
60.47
96.25
98.14
500
69.44
65.85
59.23
95.88
97.69
Appendix
Table 8: Generator objective, evaluated at two checkpoints of the same 100-step run.
University of Science and Technology of China, Hefei, China · Institute of Artificial Intelligence, China Telecom (TeleAI), Shanghai, China · Shanghai Jiao Tong University, Shanghai, China