Video Generation Models: A Survey of Post-Training and Alignment
Authors: Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, +5 more
Organizations: Arizona State University · Twitch · Stanford University · eBay · NewsBreak · Microsoft · Columbia University · University of Southern California · Carnegie Mellon University
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal properties. These challenges motivate systematic post-training strategies that adapt pretrained models without retraining them from scratch. In this survey, we present the first comprehensive review of post-training and alignment in video generation models. We frame post-training as a unifying framework and distinguish between implicit alignment and explicit alignment based on how alignment signals are enforced. From this perspective, we organize existing approaches into four broad categories: supervised fine-tuning methods, self-training and distillation methods, preference- and reward-based methods, and inference-time methods. This taxonomy provides a coherent view of how alignment signals shape model behavior across both training and deployment. Beyond methodological advances, we review commonly used datasets, benchmarks, and evaluation practices, and discuss open challenges such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation. This survey aims to provide a structured conceptual foundation and practical guidance for advancing controllable and reliable video generation models.
Figures & tables
Figure 1 : An overview of the post-training and alignment in video generation, and the scope of this survey.
Figure 2 : Research trends in post-training and alignment for video generation models (2022–Feb. 2026).
Figure 4 : Overview of video generation tasks . Left : Text-to-Video (T2V) models learn a video prior to map noise to pixels guided by text prompts. Middle : Image-to-Video (I2V) injects dynamics into a static image, typically freezing spatial layers and training temporal attention modules. Right : Video-to-Video (V2V) focuses on structure-preserving editing, injecting spatial guidance to align with the source layout.
Figure 5 : Architecture of a modern Latent Diffusion Transformer (DiT) for video generation. The pipeline has three stages: (1) Compression: A spatio-temporal latent encoder compresses the input video into latent representations. (2) Diffusion Modeling: Stochastic perturbations are applied in latent space, and the resulting spatio-temporal tokens ( blue tokens ) are processed by a DiT backbone together with text embeddings ( green tokens ) to learn a denoising objective. (3) Decoding: The predicted clean latents are mapped back to pixel space by a spatio-temporal latent decoder.
Model
Sub-category
Stages
Base Model
GPU
Venue
Year
Link
Tune-A-Video [ 26 ]
Instruction-followed Fine-tuning
1
Stable Diffusion
A100
ICCV
2023
VideoComposer [ 30 ]
Multi-conditional Control
2
Stable Diffusion
-
NeurIPS
2023
DreamPose [ 208 ]
Multi-conditional Control
2
Stable Diffusion
2xA100
ICCV
2023
PYoCo [ 110 ]
Domain Adaptation
4
eDiff-I
-
ICCV
2023
–
SparseCtrl [ 209 ]
Multi-conditional Control
1
Stable Diffusion
-
ECCV
2024
VideoDirectorGPT [ 210 ]
Instruction-followed Fine-tuning
1
ModelScopeT2V
8xA6000
COLM
2024
Table 1 : Summary of supervised fine-tuning methods for video generation. Entries are sorted chronologically by publication date. “-” indicates that the corresponding information is not reported or not clearly specified in the original paper.
Model
Sub-category
Stages
Base Model
GPU
Venue
Year
Link
SF-V [ 228 ]
Knowledge Distillation
1
Stable Video Diffusion
8xA100
NeurIPS
2024
CausVid [ 36 ]
Knowledge Distillation
2
Wan2.1
-
CVPR
2025
CustomTTT [ 35 ]
Self-training and Test-time Training
3
CogVideoX
A6000
AAAI
2025
DOLLAR [ 231 ]
Knowledge Distillation
3
DiT OpenSora LDM
8xA100
ICCV
2025
TDM [ 226 ]
Knowledge Distillation
1
Stable Diffusion
-
arXiv
2025
Reangle-A-Video [ 145 ]
Self-training and Test-time Training
2
CogVideoX
-
ICCV
2025
Table 2 : Summary of self-training and knowledge distillation methods for video generation. Entries are sorted chronologically by publication date. “-” indicates that the corresponding information is not reported or not clearly specified in the original paper.
Model
Sub-category
Stages
Base Model
GPU
Venue
Year
Link
InstructVideo [ 19 ]
Preference-based Optimization
1
ModelScopeT2V
4xA100
CVPR
2024
T2V-Turbo [ 281 ]
Preference-based Optimization
1
VideoCrafter2 ModelScopeT2V
8xA100
NeurIPS
2024
VADER [ 282 ]
Preference-based Optimization
1
VideoCrafter OpenSora ModelScopeT2V Stable Video Diffusion
2xA6000
arXiv
2024
Prompt-A-Video [ 33 ]
Preference-based Optimization
2
OpenSora CogVideoX
-
ICCV
2024
PersonalVideo [ 279 ]
Video Reward Modeling
1
HunyuanVideo AnimateDiff
A800
ICCV
2025
VideoDPO [ 27 ]
Preference-based Optimization
–
VideoCrafter2 T2V-Turbo CogVideo
4xA100
CVPR
2025
Table 3 : Summary of preference-based and reinforcement learning methods for video generation. Entries are sorted chronologically by publication date. “-” indicates that the corresponding information is not reported or not clearly specified in the original paper.
Model
Sub-Category
Base Model
Venue
Year
Link
Gen-1 [ 292 ]
Direct Trajectory Guidance
Stable Diffusion
ICCV
2023
MotionAgent [ 301 ]
Iterative Refinement and Self-Correction
Stable Video Diffusion
ICCV
2025
AICL [ 37 ]
Direct Trajectory Guidance
VideoCrafter VideoCrafter2 LVDM
ACM MM
2025
–
PAHA [ 293 ]
Direct Trajectory Guidance
VLDM
arXiv
2025
–
InstanceV [ 38 ]
Direct Trajectory Guidance
Wan
arXiv
2025
–
ALG [ 295 ]
Direct Trajectory Guidance
CogVideoX Wan 2.1 HunyuanVideo LTX
arXiv
2025
Table 4 : Summary of inference-time methods via post-trained signals for video generation.
Name
Size
Tasks
Link
ChronoMagic-Pro [ 312 ]
460,000
High resolution time-lapse video.
SafeSora [ 20 ]
57,333
Human preference text-video pairs for safety and value alignment.
CookGen [ 313 ]
200,000
Long-form narrative generation in the cooking domain.
HOIGen-1M [ 314 ]
1,000,000
Human-object interaction videos.
TIP-I2V [ 315 ]
1,700,000
User-driven text-image prompt dataset for image-to-video generation.
SynFMC [ 316 ]
62,000
Camera-object motion control for video generation.
Table 5 : Datasets used for training in video generation post-training and alignment.
Name
Size
Tasks
Link
FETV [ 357 ]
618
Fine-grained and temporal-aware evaluation of text-to-video generation.
StoryBench [ 359 ]
6,000
Story-driven text-to-video generation evaluation
–
VBench [ 17 ]
–
Multi-dimensional video generation evaluation.
ChronoMagic-Bench [ 312 ]
1,649
Time-lapse T2V generation; temporal coherence and metamorphic change evaluation.
EvalCrafter [ 347 ]
700
Text-to-video generation across diverse prompt types and multi-dimensional quality criteria.
TAVGBench [ 338 ]
1,700,000
Text to Audible-Video Generation.
Table 6 : Representative benchmarks used for video generation post-training and alignment evaluation.
Method
Backbone
VBench
Model
Size
TF
AQ
SC
IQ
MS
BC
DD
Supervised Fine-Tuning Methods
ReCapture [ 216 ]
Stable Video Diffusion
–
91.1
57.4
88.5
64.8
98.2
92.0
49.0
Phantom [ 200 ]
MMDiT
–
–
58.0
–
70.6
99.3
–
–
Follow-Your-Creation [ 146 ]
Wan2.1
–
88.2
–
90.3
–
92.4
89.3
–
MinT [ 213 ]
OpenSora
–
–
54.4
90.0
60.9
98.8
95.0
71.1
Table 7 : Reported quantitative results of representative post-training methods on VBench. TF: Temporal Flickering, AQ: Aesthetic Quality, SC: Subject Consistency, IQ: Imaging Quality, MS: Motion Smoothness, BC: Background Consistency, DD: Dynamic Degree. Bold values indicate the best reported score within each post-training family. “-” indicates that the corresponding information is not reported in the original paper.
Method
Backbone
VBench2
Model
Size
Creativity
Commonsense
Controllability
Human Fidelity
Physics
Supervised Fine-Tuning Methods
MoAlign [ 106 ]
CogVideoX
2B
52.8
65.5
25.7
86.7
48.8
Preference- and Reward-Based Methods
Euphonium [ 245 ]
HunyuanVideo
13B
41.4
67.2
26.9
88.9
46.8
PhysCorr [ 259 ]
Wan2.1
14B
29.6
42.3
35.4
85.1
49.1
Table 8 : Reported quantitative results of representative post-training methods on VBench2. Bold values indicate the best reported score within each post-training family. “-” indicates that the corresponding information is not reported in the original paper.
Method
Backbone
VideoPhy
VideoPhy2
Model
Size
Semantic Adherence
Physics Consistency
Semantic Adherence
Physics Consistency
Supervised Fine-Tuning Methods
MoAlign [ 106 ]
CogVideoX
2B
49.3
39.4
28.8
75.0
VideoREPA [ 107 ]
CogVideoX
5B
72.1
40.1
21.0
72.5
WISA [ 105 ]
CogVideoX
5B
67.0
38.0
–
–
Preference- and Reward-Based Methods
Table 9 : Reported quantitative results on physical-plausibility benchmarks. Bold values indicate the best reported score within each post-training family. “-” indicates that the corresponding information is not reported in the original paper.