cs.CVJul 30, 2026

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

Authors: Junhao ChenMingjin ChenHenghaofan ZhangMinglin ChenLiaoyuan FanBoran ZhangSaining ZhangMingze Sun+4 more

Organizations: Tsinghua University, China · The Hong Kong Polytechnic University, China · University of Electronic Science and Technology of China, China · Sun Yat-sen University, China · The University of Hong Kong, China · University of Science and Technology of China, China · Nanyang Technological University, Singapore · SparcAI Inc., USA

Abstract

Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D generative rendering setting raises a representation question: what image-format condition lets a video backbone obey both camera motion and scene-internal animation? We propose DAR, a reference-guided renderer that extends Wan2.2 camera control from Plücker rays alone to a joint camera-plus-geometry interface. DAR projects a neural 4D G-buffer (tracking, world position, and normal) from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior. The central design choice is the pair of tracking and world position. Tracking identifies the persistent surface element that should carry appearance; world position gives its current scene-coordinate state; normal supplies local shape. Depth plus calibrated rays can recover 3D in principle, but depth is a camera-dependent chart in which camera and object motion are mixed. On the 68-case DAR-4D benchmark, LoRA DAR reaches PSNR 23.22, SSIM 0.895, and LPIPS 0.134, improving over off-the-shelf Wan2.2-Depth by 1.54 dB PSNR; a full fine-tune reaches PSNR 25.36 and SSIM 0.917. Matched ablations show that replacing world position by depth reduces PSNR by 1.26--1.55 dB at every checkpoint, supporting tracking+world-position correspondence as a practical 4D rendering condition.

Explore similar work

CardsList
  1. Beyond Pixels: From Video Priors to 4D Worlds

    Aug 11, 2026Zihao Liu, Xiaolong Shen, Zhenglin Zhou +24D GenerationVideo Latents

  2. Feed-forward Motion In-betweening for Any 4D

    Jun 20, 2026Hiroki Nishizawa, Hubert P. H. Shum, Yoshihiro Fukuhara +24D GenerationMeshes