Video Generation Models: A Survey of Post-Training and Alignment
Authors: Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, +5 more
Organizations: Arizona State University · Twitch · Stanford University · eBay · NewsBreak · Microsoft · Columbia University · University of Southern California · Carnegie Mellon University
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal properties. These challenges motivate systematic post-training strategies that adapt pretrained models without retraining them from scratch. In this survey, we present the first comprehensive review of post-training and alignment in video generation models. We frame post-training as a unifying framework and distinguish between implicit alignment and explicit alignment based on how alignment signals are enforced. From this perspective, we organize existing approaches into four broad categories: supervised fine-tuning methods, self-training and distillation methods, preference- and reward-based methods, and inference-time methods. This taxonomy provides a coherent view of how alignment signals shape model behavior across both training and deployment. Beyond methodological advances, we review commonly used datasets, benchmarks, and evaluation practices, and discuss open challenges such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation. This survey aims to provide a structured conceptual foundation and practical guidance for advancing controllable and reliable video generation models.
Figures & tables
Figure 1 : An overview of the post-training and alignment in video generation, and the scope of this survey.
Figure 2 : Research trends in post-training and alignment for video generation models (2022–Feb. 2026).
Figure 4 : Overview of video generation tasks . Left : Text-to-Video (T2V) models learn a video prior to map noise to pixels guided by text prompts. Middle : Image-to-Video (I2V) injects dynamics into a static image, typically freezing spatial layers and training temporal attention modules. Right : Video-to-Video (V2V) focuses on structure-preserving editing, injecting spatial guidance to align with the source layout.
Figure 5 : Architecture of a modern Latent Diffusion Transformer (DiT) for video generation. The pipeline has three stages: (1) Compression: A spatio-temporal latent encoder compresses the input video into latent representations. (2) Diffusion Modeling: Stochastic perturbations are applied in latent space, and the resulting spatio-temporal tokens ( blue tokens ) are processed by a DiT backbone together with text embeddings ( green tokens ) to learn a denoising objective. (3) Decoding: The predicted clean latents are mapped back to pixel space by a spatio-temporal latent decoder.
Model
Sub-category
Stages
Base Model
GPU
Venue
Year
Link
Tune-A-Video [ 26 ]
Instruction-followed Fine-tuning
1
Stable Diffusion
A100
ICCV
2023
VideoComposer [ 30 ]
Multi-conditional Control
2
Stable Diffusion
-
NeurIPS
2023
DreamPose [ 208 ]
Multi-conditional Control
2
Stable Diffusion
2xA100
ICCV
2023
PYoCo [ 110 ]
Domain Adaptation
4
eDiff-I
-
ICCV
2023
–
SparseCtrl [ 209 ]
Multi-conditional Control
1
Stable Diffusion
-
ECCV
2024
VideoDirectorGPT [ 210 ]
Instruction-followed Fine-tuning
1
ModelScopeT2V
8xA6000
COLM
2024
Table 1 : Summary of supervised fine-tuning methods for video generation. Entries are sorted chronologically by publication date. “-” indicates that the corresponding information is not reported or not clearly specified in the original paper.
Model
Sub-category
Stages
Base Model
GPU
Venue
Year
Link
SF-V [ 228 ]
Knowledge Distillation
1
Stable Video Diffusion
8xA100
NeurIPS
2024
CausVid [ 36 ]
Knowledge Distillation
2
Wan2.1
-
CVPR
2025
CustomTTT [ 35 ]
Self-training and Test-time Training
3
CogVideoX
A6000
AAAI
2025
DOLLAR [ 231 ]
Knowledge Distillation
3
DiT OpenSora LDM
8xA100
ICCV
2025
TDM [ 226 ]
Knowledge Distillation
1
Stable Diffusion
-
arXiv
2025
Reangle-A-Video [ 145 ]
Self-training and Test-time Training
2
CogVideoX
-
ICCV
2025
Table 2 : Summary of self-training and knowledge distillation methods for video generation. Entries are sorted chronologically by publication date. “-” indicates that the corresponding information is not reported or not clearly specified in the original paper.
Model
Sub-category
Stages
Base Model
GPU
Venue
Year
Link
InstructVideo [ 19 ]
Preference-based Optimization
1
ModelScopeT2V
4xA100
CVPR
2024
T2V-Turbo [ 281 ]
Preference-based Optimization
1
VideoCrafter2 ModelScopeT2V
8xA100
NeurIPS
2024
VADER [ 282 ]
Preference-based Optimization
1
VideoCrafter OpenSora ModelScopeT2V Stable Video Diffusion
2xA6000
arXiv
2024
Prompt-A-Video [ 33 ]
Preference-based Optimization
2
OpenSora CogVideoX
-
ICCV
2024
PersonalVideo [ 279 ]
Video Reward Modeling
1
HunyuanVideo AnimateDiff
A800
ICCV
2025
VideoDPO [ 27 ]
Preference-based Optimization
–
VideoCrafter2 T2V-Turbo CogVideo
4xA100
CVPR
2025
Table 3 : Summary of preference-based and reinforcement learning methods for video generation. Entries are sorted chronologically by publication date. “-” indicates that the corresponding information is not reported or not clearly specified in the original paper.
Model
Sub-Category
Base Model
Venue
Year
Link
Gen-1 [ 292 ]
Direct Trajectory Guidance
Stable Diffusion
ICCV
2023
MotionAgent [ 301 ]
Iterative Refinement and Self-Correction
Stable Video Diffusion
ICCV
2025
AICL [ 37 ]
Direct Trajectory Guidance
VideoCrafter VideoCrafter2 LVDM
ACM MM
2025
–
PAHA [ 293 ]
Direct Trajectory Guidance
VLDM
arXiv
2025
–
InstanceV [ 38 ]
Direct Trajectory Guidance
Wan
arXiv
2025
–
ALG [ 295 ]
Direct Trajectory Guidance
CogVideoX Wan 2.1 HunyuanVideo LTX
arXiv
2025
Table 4 : Summary of inference-time methods via post-trained signals for video generation.
Name
Size
Tasks
Link
ChronoMagic-Pro [ 312 ]
460,000
High resolution time-lapse video.
SafeSora [ 20 ]
57,333
Human preference text-video pairs for safety and value alignment.
CookGen [ 313 ]
200,000
Long-form narrative generation in the cooking domain.
HOIGen-1M [ 314 ]
1,000,000
Human-object interaction videos.
TIP-I2V [ 315 ]
1,700,000
User-driven text-image prompt dataset for image-to-video generation.
SynFMC [ 316 ]
62,000
Camera-object motion control for video generation.
Table 5 : Datasets used for training in video generation post-training and alignment.
Name
Size
Tasks
Link
FETV [ 357 ]
618
Fine-grained and temporal-aware evaluation of text-to-video generation.
StoryBench [ 359 ]
6,000
Story-driven text-to-video generation evaluation
–
VBench [ 17 ]
–
Multi-dimensional video generation evaluation.
ChronoMagic-Bench [ 312 ]
1,649
Time-lapse T2V generation; temporal coherence and metamorphic change evaluation.
EvalCrafter [ 347 ]
700
Text-to-video generation across diverse prompt types and multi-dimensional quality criteria.
TAVGBench [ 338 ]
1,700,000
Text to Audible-Video Generation.
Table 6 : Representative benchmarks used for video generation post-training and alignment evaluation.
Method
Backbone
VBench
Model
Size
TF
AQ
SC
IQ
MS
BC
DD
Supervised Fine-Tuning Methods
ReCapture [ 216 ]
Stable Video Diffusion
–
91.1
57.4
88.5
64.8
98.2
92.0
49.0
Phantom [ 200 ]
MMDiT
–
–
58.0
–
70.6
99.3
–
–
Follow-Your-Creation [ 146 ]
Wan2.1
–
88.2
–
90.3
–
92.4
89.3
–
MinT [ 213 ]
OpenSora
–
–
54.4
90.0
60.9
98.8
95.0
71.1
Table 7 : Reported quantitative results of representative post-training methods on VBench. TF: Temporal Flickering, AQ: Aesthetic Quality, SC: Subject Consistency, IQ: Imaging Quality, MS: Motion Smoothness, BC: Background Consistency, DD: Dynamic Degree. Bold values indicate the best reported score within each post-training family. “-” indicates that the corresponding information is not reported in the original paper.
Method
Backbone
VBench2
Model
Size
Creativity
Commonsense
Controllability
Human Fidelity
Physics
Supervised Fine-Tuning Methods
MoAlign [ 106 ]
CogVideoX
2B
52.8
65.5
25.7
86.7
48.8
Preference- and Reward-Based Methods
Euphonium [ 245 ]
HunyuanVideo
13B
41.4
67.2
26.9
88.9
46.8
PhysCorr [ 259 ]
Wan2.1
14B
29.6
42.3
35.4
85.1
49.1
Table 8 : Reported quantitative results of representative post-training methods on VBench2. Bold values indicate the best reported score within each post-training family. “-” indicates that the corresponding information is not reported in the original paper.
Method
Backbone
VideoPhy
VideoPhy2
Model
Size
Semantic Adherence
Physics Consistency
Semantic Adherence
Physics Consistency
Supervised Fine-Tuning Methods
MoAlign [ 106 ]
CogVideoX
2B
49.3
39.4
28.8
75.0
VideoREPA [ 107 ]
CogVideoX
5B
72.1
40.1
21.0
72.5
WISA [ 105 ]
CogVideoX
5B
67.0
38.0
–
–
Preference- and Reward-Based Methods
Table 9 : Reported quantitative results on physical-plausibility benchmarks. Bold values indicate the best reported score within each post-training family. “-” indicates that the corresponding information is not reported in the original paper.
While large-scale video diffusion models have demonstrated impressive capabilities in generating high-resolution and semantically rich content, a significant gap remains between their pretraining performance and real-world deployment requirements due to critical issues such as prompt sensitivity, temporal inconsistency, and prohibitive inference costs. To bridge this gap, we propose a comprehensive post-training framework that systematically aligns pretrained models with user intentions through four synergistic stages: we first employ Supervised Fine-Tuning (SFT) to transform the base model into a stable instruction-following policy, followed by a Reinforcement Learning from Human Feedback (RLHF) stage that utilizes a novel Group Relative Policy Optimization (GRPO) method tailored for video diffusion to enhance perceptual quality and temporal coherence; subsequently, we integrate Prompt Enhancement via a specialized language model to refine user inputs, and finally address system efficiency through Inference Optimization. Together, these components provide a systematic approach to improving visual quality, temporal coherence, and instruction following, while preserving the controllability learned during pretraining. The result is a practical blueprint for building scalable post-training pipelines that are stable, adaptable, and effective in real-world deployment. Extensive experiments demonstrate that this unified pipeline effectively mitigates common artifacts and significantly improves controllability and visual aesthetics while adhering to strict sampling cost constraints.
Zeyue Xue, Siming Fu, Jie Huang +9
The University of Hong Kong · JD Explore Academy · Tsinghua University +2
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.
Yuandong Pu, Le Zhuo, Sayak Paul +11
Shanghai Jiao Tong University · Shanghai AI Laboratory · Krea AI +4
Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework that scales to modern, large, high-quality video datasets and models. We use this to study which fine-tuning clips improve or degrade temporal dynamics. Motive isolates temporal dynamics from static appearance via motion-weighted loss masks, yielding efficient and scalable motion-specific influence computation. On text-to-video models, Motive identifies clips that strongly affect motion and guides data curation that improves temporal consistency and physical plausibility. With Motive-selected high-influence data, our method improves both motion smoothness and dynamic degree on VBench, achieving a 74.1% human preference win rate compared with the pretrained base model. To our knowledge, this is the first framework to attribute motion rather than visual appearance in video generative models and to use it to curate fine-tuning data.
Xindi Wu, Despoina Paschalidou, Jun Gao +5
NVIDIA · Princeton University · University of Michigan +3