cs.CVMar 25, 2026

OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning

Authors: Kaihang Pan, Qi Tian, Jianwei Zhang, Weijie Kong, Jiangfeng Xiong, Yanxin Long, Shixue Zhang, Haiyi Qiu, +7 more

Organizations: Zhejiang University · Tencent Hunyuan · Nanyang Technological University

Abstract

While proprietary systems such as Seedance-2.0 have achieved remarkable success in omni-capable video generation, the academic research community lags far behind: most of its models remain heavily fragmented, and the few existing efforts toward unified video generation still struggle to seamlessly integrate diverse tasks within a single framework. To bridge this gap, we propose OmniWeaving, an omni-level video generation model featuring powerful multimodal composition and reasoning-informed capabilities. By leveraging a massive-scale pretraining dataset that encompasses diverse compositional and reasoning-augmented scenarios, OmniWeaving learns to temporally bind interleaved text, multi-image, and video inputs while acting as an intelligent agent to infer complex user intentions for sophisticated video creation. Furthermore, we introduce IntelligentVBench, the first comprehensive benchmark designed to rigorously assess next-level intelligent unified video generation. Extensive experiments demonstrate that OmniWeaving achieves SoTA performance among open-source academic unified models. The code and model are publicly available. Project Page: https://omniweaving.github.io.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. UniVideo: Unified Understanding, Generation, and Editing for Videos

    Oct 9, 2025Cong Wei, Quande Liu, Zixuan Ye +5Video EditingMultimodal Generation

  2. Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models

    May 29, 2026Jiazheng Xing, Hangjie Yuan, Lingling Cai +9Generative Video ModelsVisual Fidelity