We argue that as video generation extends to longer durations, subject consistency should be evaluated over \textit{dynamic subject sets}. We therefore introduce \textbf{DynSC-Eval}, an evaluation framework that dynamically tracks eligible subjects throughout their visible lifespans and measures local continuity and global identity preservation using six complementary object-level metrics, with explicit detection of inconsistency events. To validate its effectiveness, we design synthetic experiments that actively inject inconsistency events, demonstrating both the sensitivity of DynSC-Eval and the limitations of existing metrics. Evaluations of diverse models on 5s, 15s, and 60s video generation further reveal substantial subject consistency differences that are obscured by conventional metrics. Beyond evaluation, we construct rewards from DynSC-Eval and apply DiffusionNFT post-training in an autonomous-driving testbed. On 5s generation, our approach reduces the six inconsistency metrics by an average of 13.82% for Wan-2.1-1.3B and 5.66% for SANA-2B, with improvements also observed on the I2V model ReSim. Qualitative comparisons further demonstrate the effectiveness of our method. We then extend generation to 10s and 30s through curriculum learning and show that consistency optimization remains effective while largely preserving other capabilities.
Figures & tables
Figure 1: Examples of subject inconsistency in long video generation. As the generation duration increases, subject inconsistency events become increasingly prevalent. The corresponding prompts are provided in Appendix A.7 . Additional results are presented on the project website .
Figure 2: Overview of DynSC-Eval. The framework dynamically tracks subjects and their lifecycle, detects inconsistency events (e.g., sudden appearance/disappearance), and decomposes consistency into local and global consistency, enabling fine-grained evaluation of dynamic subject sets.
Dimension
Metric
Measured Property
Local
Spike Rate (SR)
Fraction of subjects with inconsistency events
Local
Events/100 Track-s (ER)
Frequency of inconsistency events
Local
Mean Adjacent Drift (MAD)
Track-averaged mean adjacent appearance distance
Global
Endpoint Drift (ED)
Start-to-end identity shift
Global
Mean Reference Drift (MRD)
Average long-term identity drift
Global
Worst-frame Drift (WFD)
Maximum identity deviation
Table 1: Summary of DynSC-Eval metrics.
Figure 3: Controlled synthetic experiments with actively injected inconsistencies. We inject local inconsistencies (shape mutation, sudden appearance, sudden disappearance) and global inconsistencies (smooth identity drift) to evaluate metric sensitivity. As the corruption rate increases, our metrics (SR and ED) show a clear response, while VBench (1-SC) largely fails to capture these failures.
DynSC-Eval Metrics
VBench Metrics
Local Consistency
Global Consistency
Model
SR ↓
ER ↓
MAD ↓
ED ↓
MRD ↓
WFD ↓
SC ↑
Dyn ↑
QS ↑
Open-source Models (5s)
LTX-Video-2B ( HaCohen et al., 2024 )
0.3628
28.39
0.2629
0.3665
0.3475
0.5474
84.58
96.00
73.82
Wan-2.1-1.3B ( Wang et al., 2025 )
0.2530
10.96
0.1757
0.3384
0.2717
0.5041
88.09
96.00
77.09
SANA-2B ( Chen et al., 2025b )
0.1932
7.214
0.1224
0.3035
0.2081
0.4108
94.10
99.00
82.75
Table 2: Subject consistency evaluation across diverse models. DynSC-Eval metrics are reported to four significant figures (lower is better). VBench scores are reported on a 0–100 scale (higher is better): SC denotes Subject Consistency, Dyn denotes Dynamic Degree, QS denotes the Quality Score aggregated over seven quality dimensions. Best results within each group are in bold.
Figure 4: Overview of consistency reward optimization framework. Given text or reference-image conditioning, the video model generates videos that DynSC-Eval assesses through object-level local and global consistency metrics. These metrics are combined into a consistency reward for DiffusionNFT optimization, forming a feedback loop to improve temporal consistency. The framework applies to both T2V and I2V generation across short and long durations.
Figure 5: Qualitative comparison on Wan-2.1-1.3B. We compare the baseline with our method after DiffusionNFT. Top: our method eliminates the drift of a distant car in later frames. Middle: the parked car on the left maintains a more stable shape. Bottom: the interpenetration and deformation artifact in the fifth snapshot is removed. Overall, our method produces visually more stable and temporally consistent videos. Additional qualitative results are provided on the project website .
DynSC-Eval Metrics
VBench Metrics
Local Consistency
Global Consistency
Model
Setting
SR ↓
ER ↓
MAD ↓
ED ↓
MRD ↓
WFD ↓
Mot ↑
Dyn ↑
Aes ↑
Img ↑
Text-to-Video (T2V)
Wan-2.1-1.3B ( Wang et al., 2025 )
Baseline
0.2560
9.177
0.1059
0.3837
0.2393
0.4428
98.35
92.00
50.47
71.40
+DiffNFT
0.2220
7.139
0.08829
0.3446
0.2129
0.4002
98.49
92.00
50.81
71.82
SANA-2B ( Chen et al., 2025b )
Baseline
0.2835
10.50
0.1313
0.3979
0.2733
0.4734
98.84
91.00
53.83
71.92
Table 3: Results of T2V and I2V models. +DiffNFT denotes the results after 75 optimization steps. DynSC-Eval metrics are reported to four significant figures (lower is better). VBench scores use a 0–100 scale (higher is better). Better DynSC-Eval results are highlighted in bold.
DynSC-Eval Metrics
VBench Metrics
Local Consistency
Global Consistency
Setting
SR ↓
ER ↓
MAD ↓
ED ↓
MRD ↓
WFD ↓
Mot ↑
Dyn ↑
Aes ↑
Img ↑
10s baseline
0.3843
11.77
0.1328
0.4927
0.3291
0.5759
98.03
98.00
50.49
72.12
+DiffNFT
0.3350
9.172
0.1119
0.4754
0.3087
0.5547
97.96
98.00
50.04
72.86
30s baseline
0.4684
13.27
0.1675
0.5588
0.3712
0.6435
96.48
98.00
48.45
67.92
+DiffNFT
0.4473
12.00
0.1633
0.5088
0.3373
0.5973
96.89
98.00
48.85
67.45
Table 4: Results of the 10s and 30s models. +DiffNFT denotes the results after 100 optimization steps. DynSC-Eval metrics are reported to four significant figures (lower is better). VBench scores use a 0–100 scale (higher is better). Better DynSC-Eval results are highlighted in bold.
DynSC-Eval Metrics
VBench Metrics
Local Consistency
Global Consistency
Reward
SR ↓
ER ↓
MAD ↓
ED ↓
MRD ↓
WFD ↓
Mot ↑
Dyn ↑
Aes ↑
Img ↑
5s baseline
0.2560
9.177
0.1059
0.3837
0.2393
0.4428
98.35
92.00
50.47
71.40
Local-only
0.2443
7.549
0.08653
0.3434
0.2110
0.3929
98.61
85.00
49.81
70.84
Global-only
0.2691
9.155
0.09702
0.3583
0.2262
0.4205
98.23
91.00
50.59
73.12
Ours (local+global)
0.2220
7.139
0.08829
0.3446
0.2129
0.4002
98.49
92.00
50.81
71.82
Table 5: Ablation of consistency reward components. All optimized models are evaluated at Step 75 using the same configuration. DynSC-Eval metrics are reported to four significant figures (lower is better). VBench scores use a 0–100 scale (higher is better).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Metric responses to Local-only inconsistency injection. As the fraction of objects with local inconsistencies increases, SR and ER rise, while ED remains nearly unchanged. MRD also responds to these temporary deviations from the reference appearance. Each point shows the mean over three random seeds, with error bars indicating one sample standard deviation. VBench is shown as 1−SC on the right axis.
Figure 7: Metric responses to Global-only inconsistency injection. As the fraction of objects undergoing smooth shape transitions increases, ED and MRD rise, while SR and ER remain so close to clean baseline and change minimally. Each point shows the mean over three random seeds, with error bars indicating one sample standard deviation. VBench is shown as 1−SC on the right axis.
Figure 8: Metric responses to joint local and global inconsistency injection. As the fraction of perturbed objects increases, SR closely follows the ideal relation SR=ρ , while ER, ED, and MRD increase. Each point shows the mean over three random seeds, with error bars indicating one sample standard deviation. VBench is shown as 1−SC on the right axis.
Setting
5s
10s
30s
Training clips
68,400
34,208
24,016
Validation clips
3,600
1,792
48
Optimizer
AdamW
Adam betas
(0.9,0.999)
Adam epsilon
10−8
Appendix
Table 6: Settings for 5s, 10s, and 30s fine-tuning.
Figure 9: Validation loss during 5s Wan-2.1-1.3B fine-tuning.
Figure 10: Validation loss during 5s SANA-2B fine-tuning.
Figure 11: (a). Validation loss during 10s fine-tuning. (b). Validation loss during 30s fine-tuning.
DynSC-Eval Metrics
VBench Metrics
Local Consistency
Global Consistency
Setting
SR ↓
ER ↓
MAD ↓
ED ↓
MRD ↓
WFD ↓
Mot ↑
Dyn ↑
Aes ↑
Img ↑
5s baseline
0.2560
9.177
0.1059
0.3837
0.2393
0.4428
98.35
92.00
50.47
71.40
w/o Reg.
0.2397
7.853
0.09514
0.3514
0.2223
0.4096
98.36
91.00
51.10
73.13
w/ Reg.
0.2220
7.139
0.08829
0.3446
0.2129
0.4002
98.49
92.00
50.81
71.82
Appendix
Table 7: Ablation on regularization for Wan-2.1-1.3B on 5s generation. Both optimized variants use the same combined consistency reward and are evaluated at Step 75. Regularization consistently improves DynSC-Eval metrics while preserving motion dynamics. DynSC-Eval metrics use four significant figures, and VBench scores use a 0–100 scale.
Figure 12: Effect of reference regularization on video consistency during fine-tuning
Subject-driven video generation (SDV-Gen) aims to produce videos of a specific subject by adapting a pretrained video model, enabling personalized and application-driven content creation. To achieve this goal, per-subject tuning methods require approximately 200 A100 GPU hours to generate a customized video, whereas zero-shot methods avoid per-subject tuning but typically rely on millions of subject-video pairs for the supervision, incurring massive network fine-tuning costs (10K-200K A100 GPU hours). We propose a data- and compute-efficient zero-shot SDV-Gen framework that avoids test-time per-subject tuning and the use of large-scale subject-video pairs. Our key idea decomposes SDV-Gen into (i) identity injection learned from subject-image pairs and (ii) motion-awareness preservation maintained by a small set of arbitrary videos. We optimize the two tasks with stochastic switching, using random reference-frame sampling and image-token dropout to prevent trivial first-frame copying. Our gradient analysis shows that the two objectives rapidly evolve toward nearly orthogonal update subspaces, explaining the stable optimization. Using CogVideoX-5B, we adapt a single model with 200K subject-image pairs and 4,000 arbitrary videos in 288 A100 GPU hours. This yields about 1% of compute compared to prior zero-shot baselines (i.e., 0.4% of VACE and 2.8% of Phantom) while using no subject-video pairs, yet remaining competitive in subject fidelity and motion quality. We show that the same recipe transfers to Wan 2.1-1.3B and Wan 2.2-5B.
Daneul Kim, Jingxu Zhang, Wonjoon Jin +4
1Seoul National University, Seoul, Republic of Korea · 3Microsoft Research Asia, Beijing, China · 2POSTECH, Pohang, Republic of Korea
Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under camera motion. Existing solutions either improve consistency as a byproduct, apply only to static scenes or realign the latent space of the model completely. We introduce a geometry-consistency reward that directly measures whether motion in a generated video is compatible with a coherent scene. Our key insight is that in physically consistent videos, background motion should be explainable by rigid camera-induced flow, while independently moving objects should preserve appearance identity along motion trajectories. We operationalize this using optical flow, depth--pose predictions, and feature-based correspondence to separate rigid and dynamic regions and evaluate their respective consistency. Integrating this reward with reinforcement fine-tuning transforms geometric consistency from an emergent property into an explicit optimization objective for video generators. The approach is model agnostic and applies to diverse dynamic scenes containing both camera and object motion. Experiments show substantial reductions in temporal geometric artifacts over strong baselines while preserving perceptual quality. Code and model weights are published.
Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.
Zhenghao Ni, Weimin Qiu, Meng Tang
University of Toronto · University of California, Merced