While proprietary systems such as Seedance-2.0 have achieved remarkable success in omni-capable video generation, the academic research community lags far behind: most of its models remain heavily fragmented, and the few existing efforts toward unified video generation still struggle to seamlessly integrate diverse tasks within a single framework. To bridge this gap, we propose OmniWeaving, an omni-level video generation model featuring powerful multimodal composition and reasoning-informed capabilities. By leveraging a massive-scale pretraining dataset that encompasses diverse compositional and reasoning-augmented scenarios, OmniWeaving learns to temporally bind interleaved text, multi-image, and video inputs while acting as an intelligent agent to infer complex user intentions for sophisticated video creation. Furthermore, we introduce IntelligentVBench, the first comprehensive benchmark designed to rigorously assess next-level intelligent unified video generation. Extensive experiments demonstrate that OmniWeaving achieves SoTA performance among open-source academic unified models. The code and model are publicly available. Project Page: https://omniweaving.github.io.
Figures & tables
Figure 1: Showcase of OmniWeaving across diverse video generation scenarios, such as foundational tasks, multimodal composition tasks, and reasoning-augmented scenarios.
Figure 2: Training data construction pipeline for Multimodal Composition Tasks.
Figure 3: OmniWeaving consists of an MLLM for multimodal understanding and an MMDiT for generation. On this basis, we activate the thinking mode of the MLLM and further introduce the DeepStacking mechanism.
Stage1
Stage2
Stage3
MLLM
MMDiT
Hyperparameters
Learning rate
2.0×10−5
2.0×10−5
3.0×10−5
1.0×10−5
LR scheduler
Constant
Constant
Cosine
Constant
Min Learning rate
2.0×10−5
2.0×10−5
2.0×10−6
1.0×10−5
Weight decay
0.01
0.01
0.05
0.01
Table 1: Training recipe of OmniWeaving.
Benchmark
#Size
Multi-
Interleaved
Input modality
Ability to Measure
VLM-as-a-Judge
Task
Input
Image
Multi-Images
Video
Foundation
Composition
Reasoning
VBench ( Huang et al., 2024 )
946
✗
✗
✗
✗
✗
✓
✗
✗
✗
VBench++ ( Huang et al., 2025 )
2384
✓
✗
✓
✗
✗
✓
✗
✗
✗
OpenVE-Bench ( He et al., 2025 )
431
✗
✗
✗
✗
✓
✓
✓
✗
✓
TGVE+ ( Singer et al., 2024 )
1417
✗
✗
✗
✗
✓
✓
✓
✗
✗
OpenS2V-Eval ( Yuan et al., 2025 )
180
✗
✓
✓
✓
✗
✓
✓
✗
✗
Table 2: Compare IntelligentVBench with existing video generation benchmarks.
Figure 4: Examples of each task type in IntelligentVBench.
Model
#GenParams
Implicit I2V
Interpolative DI2V
TIV2V
IF↑
CP↑
VQ↑
MIN
AVG
IF↑
CP↑
VQ↑
MIN
AVG
IF↑
CP↑
VQ↑
MIN
AVG
Specialized Video Generation Models
CogVideoX-I2V
5B
3.53
3.74
2.90
2.68
3.39
-
-
-
-
-
-
-
-
-
-
Wan2.1-I2V
14B
3.76
3.33
3.34
2.78
3.48
-
-
-
-
-
-
-
-
-
-
Wan2.1-FLF2V
14B
-
-
-
-
-
4.48
4.28
4.49
3.98
4.42
-
-
-
-
-
HunyuanVideo-I2V
13B
3.37
4.17
3.87
3.00
3.80
-
-
-
-
-
-
-
-
-
-
Table 3: Main results of the Implicit I2V, Interpolative DI2V, and TIV2V tasks in IntelligentVBench. The best results are in bold for specialized- and unified- models, with the second best underlined . #GenParams reports only the generation parameters.
Model
#GenParams
1 Subject (w/ or w/o BKG)
2 Subjects (w/ or w/o BKG)
3 Subjects (w/ or w/o BKG)
IF↑
CP↑
VQ↑
MIN
AVG
IF↑
CP↑
VQ↑
MIN
AVG
IF↑
CP↑
VQ↑
MIN
AVG
Specialized Video Generation Models
SkyReels-A2
14B
3.51
4.08
4.46
3.24
4.02
3.22
3.76
4.37
2.97
3.78
1.64
1.76
2.50
1.56
1.97
SkyReels-V3
14B
3.46
3.71
4.65
2.98
3.94
3.28
3.84
4.44
3.04
3.86
2.59
3.10
4.30
2.37
3.33
MAGREF
14B
3.15
2.48
4.32
2.18
3.32
3.04
2.81
4.33
2.44
3.39
2.50
2.21
4.46
2.07
3.06
Phantom
14B
3.21
2.95
4.29
2.47
3.48
2.88
3.42
4.38
2.62
3.56
2.36
2.79
4.21
2.20
3.12
Table 4: Main results of Compositional MI2V task in IntelligentVBench. We define 3 sub-categories based on the number of subjects presented in the 1–4 input images.
Model
#GenParams
Global-Style ↑
BKG-Change ↑
Local-Change ↑
Local-RM ↑
Local-Add ↑
Subtitle-Edit ↑
Overall ↑
Specialized Video Generation Models
OmniVideo
1.3B
1.11
1.18
1.14
1.14
1.36
1.00
1.16
InsViE
2B
2.20
1.06
1.48
1.36
1.17
2.18
1.58
Lucy-Edit
5B
2.27
1.57
3.20
1.75
2.30
1.61
2.12
ICVE
13B
2.22
1.62
2.57
2.51
1.97
2.09
2.16
Ditto
14B
4.01
1.68
2.03
1.53
1.41
2.81
2.25
Table 5: Main results on OpenVE-Bench. The best results are in bold for specialized- and unified- models, with the second best among unified models underlined .
Model
#GenParams
Quality ↑
Semantic ↑
Total ↑
Specialized Video Generation Models
StepVideo
30B
84.46
71.28
81.83
CogVideoX
5B
83.05
77.33
81.91
HunyuanVideo
13B
85.09
75.82
83.24
Wan2.1
14B
85.59
76.11
83.69
Unified Video Generation Models
Table 6: Main Results on VBench.
Figure 5: (a) AVG performance when enabling or disabling the thinking mode of OmniWeaving. (b) AVG performance with different DeepStacking strategies. (c) Performance visualization across different input formats for each unified video generation model.
Figure 6: Qualitative comparison of VINO, UniVideo, and OmniWeaving.
VLMs
IF↑
CP↑
VQ↑
AVG
Qwen3-VL-235B
0.66
0.54
0.59
0.60
Seed-1.6
0.72
0.63
0.67
0.67
GPT-5
0.75
0.72
0.71
0.73
Gemini2.5-Pro
0.81
0.74
0.72
0.76
Table 7: Pearson correlation with human ratings for each VLM.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: System prompts for Text-to-Video and First-Frame-to-Video generation tasks, where <|user_img|> is the user input image and <|user_text|> is the user input prompt.
Figure 8: System prompts for Key-Frames-to-Video generation and Video-to-Video editing tasks, where <|user_text|> is the user input prompt, <|key_frame_i|> is the user input image, and <|user_video|> is the user input video.
Figure 9: System prompts for Interleaved Text-and-Multi-Image-to-Video and Text-Image-Video-to-Video generation tasks, where <|user_img_i|> is the user input images, and <|user_video|> is the user input video.
Figure 10: Two system prompt examples for reasoning-augmented video generation tasks.
Figure 11: Qualitative results for OmniWeaving on Text-to-Video, First-Frame-to-Video, and Key-Frames-to-Video generation tasks.
Figure 12: Qualitative results for OmniWeaving on Video-to-Video editing tasks.
Figure 13: Qualitative results for OmniWeaving on Compositional Multi-Image-to-Video generation tasks.
Figure 14: Qualitative results for OmniWeaving on Text-Image-Video-to-Video generation tasks.
Figure 15: Qualitative results for OmniWeaving on Reasoning-Augmented video generation.
Figure 16: Evaluation prompt template for Implicit I2V task.
Figure 17: Evaluation prompt template for Interpolative DI2V task.
Figure 18: Evaluation prompt template for Compositional MI2V task with one input image.
Figure 19: Evaluation prompt template for Compositional MI2V task with multiple input images.
Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexplored. A central challenge is that video understanding favors compact, discriminative semantic representations, whereas video generation requires dense signals that preserve visual details and temporal coherence. Videos naturally capture both spatial semantics and temporal dynamics, making them a more suitable modality for unified multimodal modeling compared to static images. In this paper, we propose Vega, a unified framework that bridges video understanding and generation. Vega leverages a shared vocabulary to jointly model text and visual representations and employs a hybrid architecture combining autoregressive (AR) prediction with diffusion-based rendering. Specifically, the AR model focuses on predicting semantically meaningful visual tokens for keyframes, providing a structured representation that guides the diffusion module in rendering dense, high-resolution video frames. Extensive experiments demonstrate that Vega achieves strong performance on video generation benchmarks such as VBench and video understanding benchmarks like VideoMME.
Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified modeling to the video domain. UniVideo adopts a dual-stream design, combining a Multimodal Large Language Model (MLLM) for instruction understanding with a Multimodal DiT (MMDiT) for video generation. This design preserves the MLLM's original text generation capabilities, enables accurate interpretation of complex multimodal instructions, and maintains visual consistency in the generated content. Built on this architecture, UniVideo unifies diverse video generation and editing tasks under a single multimodal instruction paradigm and is jointly trained across them. Extensive experiments demonstrate that UniVideo matches or surpasses state-of-the-art task-specific baselines in text/image-to-video generation, in-context video generation and in-context video editing. Notably, the unified design of UniVideo enables two forms of generalization. First, UniVideo supports task composition, such as combining editing with style transfer, by integrating multiple capabilities within a single instruction. Second, even without explicit training on free-form video editing, UniVideo transfers its editing capability from large-scale image editing data to this setting, handling unseen instructions such as changing the environment or altering materials within a video. Beyond these core capabilities, UniVideo also supports visual-prompt-based video generation, where the MLLM interprets visual prompts and guides the MMDiT during synthesis. To foster future research, we released our model and code.
Cong Wei, Quande Liu, Zixuan Ye +5
University of Waterloo · Kling Team, Kuaishou Technology
Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large high-fidelity generator into the unified training loop is computationally prohibitive, limiting achievable visual quality. We therefore propose Lumos-Nexus, a training-efficient unified video generation framework that facilitates the development of strong reasoning-driven generation capabilities while significantly enhancing visual fidelity. Lumos-Nexus adopts a two-stage design: 1) During training, only a lightweight generator is aligned with the understanding block to learn to take in reasoning-driven semantic control. 2) During inference, we introduce Unified Progressive Frequency Bridging (UPFB) to progressively hand off generation to a high-capacity pretrained generator in the shared latent space, enabling coarse-to-fine refinement and producing high-fidelity videos without compromising reasoning quality. To fill the gap in reasoning-driven video generation benchmarks, we introduce VR-Bench, which assesses a model's capability to translate inferred intent into coherent and semantically aligned video content. Extensive experiments demonstrate that Lumos-Nexus achieves substantial gains in visual realism and temporal coherence on VBench, while exhibiting strong reasoning-based generative performance on VR-Bench. Code and models are available at https://jiazheng-xing.github.io/nexus-lumos-home/.
Jiazheng Xing, Hangjie Yuan, Lingling Cai +9
Zhejiang University · National University of Singapore · DAMO Academy, Alibaba Group +4