Image-to-Video Diffusion: From Foundations to Open Frontiers
Authors: Xianlong Wang, Wenbo Pan, Shijia Zhou, Ke Li, Yuqi Wang, Zeyu Ye, Hangtao Zhang, Leo Yu Zhang, +1 more
Organizations: Department of Computer Science, City University of Hong Kong, Hong Kong SAR, China · School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan 430074, China · School of Cyber Science and Technology, University of Science and Technology of China, Hefei 230026, China · Faculty of Science and Technology, Beijing Normal University-Hong Kong Baptist University, Zhuhai 519087, China · School of Computer Science, Xiangtan University, XiangTan 411105, China · School of Information and Communication Technology, Griffith University, Southport, QLD 4215, Australia
Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation settings, this task places stricter demands on content consistency, identity preservation, and motion coherence. Although the literature grows rapidly, existing works mostly discuss I2V generation within broader topics and still lack a dedicated taxonomy together with a systematic analysis centered on this field. This work addresses that gap by treating diffusion I2V generation as a standalone subject. It first reviews the task formulation, model architectures, datasets, and evaluation metrics, and then organizes existing methods through a taxonomy based on architecture and training paradigm. It further distills four core designs, namely condition encoding, temporal modeling, noise prior design, and spatial-temporal upsampling, and discusses representative application scenarios together with major open challenges.