While proprietary systems such as Seedance-2.0 have achieved remarkable success in omni-capable video generation, the academic research community lags far behind: most of its models remain heavily fragmented, and the few existing efforts toward unified video generation still struggle to seamlessly integrate diverse tasks within a single framework. To bridge this gap, we propose OmniWeaving, an omni-level video generation model featuring powerful multimodal composition and reasoning-informed capabilities. By leveraging a massive-scale pretraining dataset that encompasses diverse compositional and reasoning-augmented scenarios, OmniWeaving learns to temporally bind interleaved text, multi-image, and video inputs while acting as an intelligent agent to infer complex user intentions for sophisticated video creation. Furthermore, we introduce IntelligentVBench, the first comprehensive benchmark designed to rigorously assess next-level intelligent unified video generation. Extensive experiments demonstrate that OmniWeaving achieves SoTA performance among open-source academic unified models. The code and model are publicly available. Project Page: https://omniweaving.github.io.
Figures & tables
Figure 1: Showcase of OmniWeaving across diverse video generation scenarios, such as foundational tasks, multimodal composition tasks, and reasoning-augmented scenarios.
Figure 2: Training data construction pipeline for Multimodal Composition Tasks.
Figure 3: OmniWeaving consists of an MLLM for multimodal understanding and an MMDiT for generation. On this basis, we activate the thinking mode of the MLLM and further introduce the DeepStacking mechanism.
Stage1
Stage2
Stage3
MLLM
MMDiT
Hyperparameters
Learning rate
2.0×10−5
2.0×10−5
3.0×10−5
1.0×10−5
LR scheduler
Constant
Constant
Cosine
Constant
Min Learning rate
2.0×10−5
2.0×10−5
2.0×10−6
1.0×10−5
Weight decay
0.01
0.01
0.05
0.01
Table 1: Training recipe of OmniWeaving.
Benchmark
#Size
Multi-
Interleaved
Input modality
Ability to Measure
VLM-as-a-Judge
Task
Input
Image
Multi-Images
Video
Foundation
Composition
Reasoning
VBench ( Huang et al., 2024 )
946
✗
✗
✗
✗
✗
✓
✗
✗
✗
VBench++ ( Huang et al., 2025 )
2384
✓
✗
✓
✗
✗
✓
✗
✗
✗
OpenVE-Bench ( He et al., 2025 )
431
✗
✗
✗
✗
✓
✓
✓
✗
✓
TGVE+ ( Singer et al., 2024 )
1417
✗
✗
✗
✗
✓
✓
✓
✗
✗
OpenS2V-Eval ( Yuan et al., 2025 )
180
✗
✓
✓
✓
✗
✓
✓
✗
✗
Table 2: Compare IntelligentVBench with existing video generation benchmarks.
Figure 4: Examples of each task type in IntelligentVBench.
Model
#GenParams
Implicit I2V
Interpolative DI2V
TIV2V
IF↑
CP↑
VQ↑
MIN
AVG
IF↑
CP↑
VQ↑
MIN
AVG
IF↑
CP↑
VQ↑
MIN
AVG
Specialized Video Generation Models
CogVideoX-I2V
5B
3.53
3.74
2.90
2.68
3.39
-
-
-
-
-
-
-
-
-
-
Wan2.1-I2V
14B
3.76
3.33
3.34
2.78
3.48
-
-
-
-
-
-
-
-
-
-
Wan2.1-FLF2V
14B
-
-
-
-
-
4.48
4.28
4.49
3.98
4.42
-
-
-
-
-
HunyuanVideo-I2V
13B
3.37
4.17
3.87
3.00
3.80
-
-
-
-
-
-
-
-
-
-
Table 3: Main results of the Implicit I2V, Interpolative DI2V, and TIV2V tasks in IntelligentVBench. The best results are in bold for specialized- and unified- models, with the second best underlined . #GenParams reports only the generation parameters.
Model
#GenParams
1 Subject (w/ or w/o BKG)
2 Subjects (w/ or w/o BKG)
3 Subjects (w/ or w/o BKG)
IF↑
CP↑
VQ↑
MIN
AVG
IF↑
CP↑
VQ↑
MIN
AVG
IF↑
CP↑
VQ↑
MIN
AVG
Specialized Video Generation Models
SkyReels-A2
14B
3.51
4.08
4.46
3.24
4.02
3.22
3.76
4.37
2.97
3.78
1.64
1.76
2.50
1.56
1.97
SkyReels-V3
14B
3.46
3.71
4.65
2.98
3.94
3.28
3.84
4.44
3.04
3.86
2.59
3.10
4.30
2.37
3.33
MAGREF
14B
3.15
2.48
4.32
2.18
3.32
3.04
2.81
4.33
2.44
3.39
2.50
2.21
4.46
2.07
3.06
Phantom
14B
3.21
2.95
4.29
2.47
3.48
2.88
3.42
4.38
2.62
3.56
2.36
2.79
4.21
2.20
3.12
Table 4: Main results of Compositional MI2V task in IntelligentVBench. We define 3 sub-categories based on the number of subjects presented in the 1–4 input images.
Model
#GenParams
Global-Style ↑
BKG-Change ↑
Local-Change ↑
Local-RM ↑
Local-Add ↑
Subtitle-Edit ↑
Overall ↑
Specialized Video Generation Models
OmniVideo
1.3B
1.11
1.18
1.14
1.14
1.36
1.00
1.16
InsViE
2B
2.20
1.06
1.48
1.36
1.17
2.18
1.58
Lucy-Edit
5B
2.27
1.57
3.20
1.75
2.30
1.61
2.12
ICVE
13B
2.22
1.62
2.57
2.51
1.97
2.09
2.16
Ditto
14B
4.01
1.68
2.03
1.53
1.41
2.81
2.25
Table 5: Main results on OpenVE-Bench. The best results are in bold for specialized- and unified- models, with the second best among unified models underlined .
Model
#GenParams
Quality ↑
Semantic ↑
Total ↑
Specialized Video Generation Models
StepVideo
30B
84.46
71.28
81.83
CogVideoX
5B
83.05
77.33
81.91
HunyuanVideo
13B
85.09
75.82
83.24
Wan2.1
14B
85.59
76.11
83.69
Unified Video Generation Models
Table 6: Main Results on VBench.
Figure 5: (a) AVG performance when enabling or disabling the thinking mode of OmniWeaving. (b) AVG performance with different DeepStacking strategies. (c) Performance visualization across different input formats for each unified video generation model.
Figure 6: Qualitative comparison of VINO, UniVideo, and OmniWeaving.
VLMs
IF↑
CP↑
VQ↑
AVG
Qwen3-VL-235B
0.66
0.54
0.59
0.60
Seed-1.6
0.72
0.63
0.67
0.67
GPT-5
0.75
0.72
0.71
0.73
Gemini2.5-Pro
0.81
0.74
0.72
0.76
Table 7: Pearson correlation with human ratings for each VLM.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: System prompts for Text-to-Video and First-Frame-to-Video generation tasks, where <|user_img|> is the user input image and <|user_text|> is the user input prompt.
Figure 8: System prompts for Key-Frames-to-Video generation and Video-to-Video editing tasks, where <|user_text|> is the user input prompt, <|key_frame_i|> is the user input image, and <|user_video|> is the user input video.
Figure 9: System prompts for Interleaved Text-and-Multi-Image-to-Video and Text-Image-Video-to-Video generation tasks, where <|user_img_i|> is the user input images, and <|user_video|> is the user input video.
Figure 10: Two system prompt examples for reasoning-augmented video generation tasks.
Figure 11: Qualitative results for OmniWeaving on Text-to-Video, First-Frame-to-Video, and Key-Frames-to-Video generation tasks.
Figure 12: Qualitative results for OmniWeaving on Video-to-Video editing tasks.
Figure 13: Qualitative results for OmniWeaving on Compositional Multi-Image-to-Video generation tasks.
Figure 14: Qualitative results for OmniWeaving on Text-Image-Video-to-Video generation tasks.
Figure 15: Qualitative results for OmniWeaving on Reasoning-Augmented video generation.
Figure 16: Evaluation prompt template for Implicit I2V task.
Figure 17: Evaluation prompt template for Interpolative DI2V task.
Figure 18: Evaluation prompt template for Compositional MI2V task with one input image.
Figure 19: Evaluation prompt template for Compositional MI2V task with multiple input images.