cs.AISep 28, 2026

CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models

Authors: Zhaolong Su, Yujin Han, Feng Wang, Jameson Dong, Hins Hu, Difan Zou

Organizations: Cornell University · The University of Hong Kong · Johns Hopkins University

Abstract

Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model's training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator's current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

    Aug 6, 2026Rui Li, Yuanzhi Liang, Ke Hao +4Reward GradientsPost-Training

  2. DRM: Diffusion-based Reward Model With Step-wise Guidance

    May 25, 2026Jaxon Zhang, Binxin Yang, Hubery Yin +2Diffusion AlignmentDiffusion Models