Towards Subject Consistency over Dynamic Subject Sets in Video Generation
Organizations: Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University
Abstract
We argue that as video generation extends to longer durations, subject consistency should be evaluated over \textit{dynamic subject sets}. We therefore introduce \textbf{DynSC-Eval}, an evaluation framework that dynamically tracks eligible subjects throughout their visible lifespans and measures local continuity and global identity preservation using six complementary object-level metrics, with explicit detection of inconsistency events. To validate its effectiveness, we design synthetic experiments that actively inject inconsistency events, demonstrating both the sensitivity of DynSC-Eval and the limitations of existing metrics. Evaluations of diverse models on 5s, 15s, and 60s video generation further reveal substantial subject consistency differences that are obscured by conventional metrics. Beyond evaluation, we construct rewards from DynSC-Eval and apply DiffusionNFT post-training in an autonomous-driving testbed. On 5s generation, our approach reduces the six inconsistency metrics by an average of 13.82% for Wan-2.1-1.3B and 5.66% for SANA-2B, with improvements also observed on the I2V model ReSim. Qualitative comparisons further demonstrate the effectiveness of our method. We then extend generation to 10s and 30s through curriculum learning and show that consistency optimization remains effective while largely preserving other capabilities.
Figures & tables
| Dimension | Metric | Measured Property |
| Local | Spike Rate (SR) | Fraction of subjects with inconsistency events |
| Local | Events/100 Track-s (ER) | Frequency of inconsistency events |
| Local | Mean Adjacent Drift (MAD) | Track-averaged mean adjacent appearance distance |
| Global | Endpoint Drift (ED) | Start-to-end identity shift |
| Global | Mean Reference Drift (MRD) | Average long-term identity drift |
| Global | Worst-frame Drift (WFD) | Maximum identity deviation |
| DynSC-Eval Metrics | VBench Metrics | ||||||||
| Local Consistency | Global Consistency | ||||||||
| Model | SR | ER | MAD | ED | MRD | WFD | SC | Dyn | QS |
| Open-source Models (5s) | |||||||||
| LTX-Video-2B ( HaCohen et al., 2024 ) | 0.3628 | 28.39 | 0.2629 | 0.3665 | 0.3475 | 0.5474 | 84.58 | 96.00 | 73.82 |
| Wan-2.1-1.3B ( Wang et al., 2025 ) | 0.2530 | 10.96 | 0.1757 | 0.3384 | 0.2717 | 0.5041 | 88.09 | 96.00 | 77.09 |
| SANA-2B ( Chen et al., 2025b ) | 0.1932 | 7.214 | 0.1224 | 0.3035 | 0.2081 | 0.4108 | 94.10 | 99.00 | 82.75 |
| DynSC-Eval Metrics | VBench Metrics | ||||||||||
| Local Consistency | Global Consistency | ||||||||||
| Model | Setting | SR | ER | MAD | ED | MRD | WFD | Mot | Dyn | Aes | Img |
| Text-to-Video (T2V) | |||||||||||
| Wan-2.1-1.3B ( Wang et al., 2025 ) | Baseline | 0.2560 | 9.177 | 0.1059 | 0.3837 | 0.2393 | 0.4428 | 98.35 | 92.00 | 50.47 | 71.40 |
| +DiffNFT | 0.2220 | 7.139 | 0.08829 | 0.3446 | 0.2129 | 0.4002 | 98.49 | 92.00 | 50.81 | 71.82 | |
| SANA-2B ( Chen et al., 2025b ) | Baseline | 0.2835 | 10.50 | 0.1313 | 0.3979 | 0.2733 | 0.4734 | 98.84 | 91.00 | 53.83 | 71.92 |
| DynSC-Eval Metrics | VBench Metrics | |||||||||
| Local Consistency | Global Consistency | |||||||||
| Setting | SR | ER | MAD | ED | MRD | WFD | Mot | Dyn | Aes | Img |
| 10s baseline | 0.3843 | 11.77 | 0.1328 | 0.4927 | 0.3291 | 0.5759 | 98.03 | 98.00 | 50.49 | 72.12 |
| +DiffNFT | 0.3350 | 9.172 | 0.1119 | 0.4754 | 0.3087 | 0.5547 | 97.96 | 98.00 | 50.04 | 72.86 |
| 30s baseline | 0.4684 | 13.27 | 0.1675 | 0.5588 | 0.3712 | 0.6435 | 96.48 | 98.00 | 48.45 | 67.92 |
| +DiffNFT | 0.4473 | 12.00 | 0.1633 | 0.5088 | 0.3373 | 0.5973 | 96.89 | 98.00 | 48.85 | 67.45 |
| DynSC-Eval Metrics | VBench Metrics | |||||||||
| Local Consistency | Global Consistency | |||||||||
| Reward | SR | ER | MAD | ED | MRD | WFD | Mot | Dyn | Aes | Img |
| 5s baseline | 0.2560 | 9.177 | 0.1059 | 0.3837 | 0.2393 | 0.4428 | 98.35 | 92.00 | 50.47 | 71.40 |
| Local-only | 0.2443 | 7.549 | 0.08653 | 0.3434 | 0.2110 | 0.3929 | 98.61 | 85.00 | 49.81 | 70.84 |
| Global-only | 0.2691 | 9.155 | 0.09702 | 0.3583 | 0.2262 | 0.4205 | 98.23 | 91.00 | 50.59 | 73.12 |
| Ours (local+global) | 0.2220 | 7.139 | 0.08829 | 0.3446 | 0.2129 | 0.4002 | 98.49 | 92.00 | 50.81 | 71.82 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | 5s | 10s | 30s |
| Training clips | 68,400 | 34,208 | 24,016 |
| Validation clips | 3,600 | 1,792 | 48 |
| Optimizer | AdamW | ||
| Adam betas | |||
| Adam epsilon | |||
| DynSC-Eval Metrics | VBench Metrics | |||||||||
| Local Consistency | Global Consistency | |||||||||
| Setting | SR | ER | MAD | ED | MRD | WFD | Mot | Dyn | Aes | Img |
| 5s baseline | 0.2560 | 9.177 | 0.1059 | 0.3837 | 0.2393 | 0.4428 | 98.35 | 92.00 | 50.47 | 71.40 |
| w/o Reg. | 0.2397 | 7.853 | 0.09514 | 0.3514 | 0.2223 | 0.4096 | 98.36 | 91.00 | 51.10 | 73.13 |
| w/ Reg. | 0.2220 | 7.139 | 0.08829 | 0.3446 | 0.2129 | 0.4002 | 98.49 | 92.00 | 50.81 | 71.82 |