cs.CVOct 1, 2026

Towards Subject Consistency over Dynamic Subject Sets in Video Generation

Authors: Tongcheng Zhang, Jun Zhu, Jianfei Chen

Organizations: Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University

Abstract

We argue that as video generation extends to longer durations, subject consistency should be evaluated over \textit{dynamic subject sets}. We therefore introduce \textbf{DynSC-Eval}, an evaluation framework that dynamically tracks eligible subjects throughout their visible lifespans and measures local continuity and global identity preservation using six complementary object-level metrics, with explicit detection of inconsistency events. To validate its effectiveness, we design synthetic experiments that actively inject inconsistency events, demonstrating both the sensitivity of DynSC-Eval and the limitations of existing metrics. Evaluations of diverse models on 5s, 15s, and 60s video generation further reveal substantial subject consistency differences that are obscured by conventional metrics. Beyond evaluation, we construct rewards from DynSC-Eval and apply DiffusionNFT post-training in an autonomous-driving testbed. On 5s generation, our approach reduces the six inconsistency metrics by an average of 13.82% for Wan-2.1-1.3B and 5.66% for SANA-2B, with improvements also observed on the I2V model ReSim. Qualitative comparisons further demonstrate the effectiveness of our method. We then extend generation to 10s and 30s through curriculum learning and show that consistency optimization remains effective while largely preserving other capabilities.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute

    Apr 23, 2025Daneul Kim, Jingxu Zhang, Wonjoon Jin +4Interactive Video GenerationVideo Generation

  2. GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation

    May 18, 2026Jan Ackermann, Shengqu Cai, Boyang Deng +3Generative Video ModelsText-To-Video Diffusion Models

  3. Learning via Self-Consistency for Diffusion-based Video Reasoning

    Sep 29, 2026Zhenghao Ni, Weimin Qiu, Meng TangVideo GenerationSelf-Consistency