cs.CVOct 8, 2026

MATE4D: Matrix-Guided Editable 4D Generation from a Single Image

Authors: Xiaotian Chen, Dongfu Yin

Organizations: Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) · Shenzhen University

Abstract

Generative models have rapidly pushed content creation be-yond 2D imagery toward dynamic 3D and 4D scene synthesis. Yet pro-ducing realistic and temporally stable 4D content from a single image is still difficult because one view provides limited structural cues and weak motion evidence. We introduce MATE4D, a framework that converts one input image into editable dynamic 4D content. Our method constructs a spatio-temporal multi-view image matrix with text-guided background manipulation, delivering coherent supervision over viewpoint, appear-ance, and motion. These synthesized observations are used to optimize 3D Gaussian primitives, which are then animated through a lightweight deformation module to form a 4D representation. The resulting scenes preserve geometry more faithfully, maintain smoother temporal behavior, and keep background edits more consistent, reducing context ambiguity and motion artifacts. Experiments on Objaverse-XL and Diffusion4D show that MATE4D outperforms strong baselines in visual quality, effi-ciency, and controllability, supporting practical AR/VR content creation.

Explore similar work

CardsList
  1. Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

    Dec 4, 2025Yanran Zhang, Ziyi Wang, Wenzhao Zheng +3Video Diffusion Models3D Scene Generation

  2. Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation

    Date pendingXiaoyan Liu, Kangrui Li, Jiaxin Liu +2Camera-Controlled Video GenerationImage-to-Video Generation

  3. Alignment Is All You Need For X-to-4D Generation

    Jul 2, 2026Qiaowei Miao, Kehan Li, Yawei Luo +1Cross-Modal AlignmentMultimodal Generation