cs.CVMar 16, 2026

Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion

Authors: Zhenghong Zhou, Xiaohang Zhan, Zhiqin Chen, Soo Ye Kim, Nanxuan Zhao, Haitian Zheng, Qing Liu, He Zhang, +3 more

Organizations: Adobe Research · University of Rochester

Abstract

Recent video diffusion models achieve strong visual quality and temporal coherence, but still lack coordinated control over scene layout, customized subject identity, and camera/subject motion. We study controllable video generation from three prompts: a first-frame scene image, 3D-aware multi-view subject references, and a motion-driving signal. This setting is challenging because background regions are often trackable under camera motion, while foreground subjects can rotate, self-occlude, and reveal new regions that require identity-consistent appearance. We introduce Tri-Prompting, a video diffusion framework that combines multi-view subject conditioning with dual-conditioned motion control. Tri-Prompting uses XYZ tracking points for visible background motion and low-resolution RGB proxies for foreground subject pose, allowing the model to recover fine appearance from multi-view references while retaining flexibility for plausible subject-scene interactions. An inference-time ControlNet scale schedule further balances motion controllability and visual realism. We evaluate Tri-Prompting against specialized baselines: DaS for motion reconstruction and Phantom for subject-driven generation. Tri-Prompting achieves competitive or improved reconstruction quality, stronger multi-view identity preservation, and better 3D consistency, while enabling 3D-aware subject insertion and in-scene manipulation.

Explore similar work

CardsList
  1. Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

    Jul 2, 2026Long Vu, Tan Ngo, Animesh Karnewar +5Video Diffusion ModelsControllability

  2. Beyond Inpainting: Unleash 3D Understanding for Stable Camera-Controlled Video Generation

    Jan 15, 2026Dong-Yu Chen, Yixin Guo, Shuojin Yang +2Camera ControlVideo Generation

  3. PE-Field 4D: Video Generation Models as Canvas

    Jul 17, 2026Yunpeng Bai, Haoxiang Li, Qixing HuangVideo GenerationCamera Trajectories