Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.
Figures & tables
Figure 1 : Composing camera and object motion in 3D from a single image. Artists choreograph foreground motion and camera paths in a shared 3D scene lifted from the input image. Each pair shows object motion with a fixed camera (top) and combined camera and object motion (bottom). The hoverboard rider follows a curved path or takes flight; the remaining pairs show a potted plant moving off a table, a swimming turtle, and a girl and dog jumping under separate controls. Each row shows the authored 3D controls followed by three generated frames. Orange denotes camera poses and paths; other colors denote foreground trajectories. Solid and dashed lines indicate elapsed and remaining motion, respectively. Best viewed as videos on the supplemental website.
Figure 2 : Why author motion in 3D? A 2D drag can leave the intended 3D rotation ambiguous. Here, the 2D-controlled baseline turns the fish in a direction inconsistent with the artist’s intent, while GenCine follows the rotation specified in 3D. Enlarged views highlight the difference.
Figure 3 : The GenCine 3D control workflow. We estimate depth and a dynamic mask from an RGB image, then lift them into a 3D scene for motion authoring. The artist specifies camera and object motion in a shared world frame using a 3D interface such as Blender. We convert the controls into background XYZ, foreground XYZ, and foreground identity maps, indicating static scene positions, moving handle positions, and handle identity. Validity masks distinguish available guidance from missing values, while the dynamic mask marks the initial foreground. A frozen Wan VAE encodes the maps for the video model. Training pairs videos with guidance recovered from their motion; at inference, the same maps express artist-authored controls.
Figure 4 : World-space XYZ maps establish correspondences across viewpoints. A moving bus is observed from different viewpoints at t=0 , t=40 , and t=80 . Static points P1 , P2 , and P3 project to different image locations but retain their world-space coordinates and encoded colors. By contrast, the highlighted region on the moving bus changes its world-space position and therefore its XYZ color over time. We consequently introduce a separate, time-invariant foreground identity map to preserve correspondence, while the foreground XYZ map represents its time-varying position.
Figure 5 : Qualitative comparison on artist-authored motion. Living room, Bus, and Robotic arm (top to bottom) each show Wan-Move, VerseCrafter, Go-with-the-Track, and GenCine (ours), in row order. The first column shows 3D controls (GenCine) or shared 2D projections (baselines); four output frames sample each control interval. Solid/dashed paths show elapsed/remaining target motion, dots mark current targets, and translucent masks show projected target regions.
Camera Shooting
Object Interactions
Camera Traj.
Quality
Consistency
Camera Traj.
Quality
Consistency
Method
Rot ↓
Trans ↓
PSNR ↑
LPIPS ↓
R-PSNR ↑
R-LPIPS ↓
Subj. ↑
Bkg. ↑
IQ ↑
Rot ↓
Trans ↓
PSNR ↑
LPIPS ↓
R-PSNR ↑
R-LPIPS ↓
Subj. ↑
Bkg. ↑
IQ ↑
Wan2.1
4.98
2.40
12.44
0.49
14.08
0.34
0.87
0.92
0.66
5.14
2.53
13.08
0.47
12.11
0.38
0.87
0.92
0.67
ATI
3.80
1.15
16.18
0.38
16.66
0.20
0.88
0.92
0.67
3.97
1.26
15.87
0.39
14.29
0.25
0.88
0.91
0.66
Wan-Move
3.73
1.02
16.58
0.36
17.13
0.19
0.90
0.93
0.69
3.76
1.08
16.51
0.37
15.46
0.22
0.89
0.94
0.67
VerseCrafter
3.92
1.27
15.27
0.41
15.82
0.27
0.90
0.92
0.69
4.02
1.24
13.94
0.43
13.27
0.33
0.89
0.92
0.68
Table 1 : Quantitative comparison on examples from Camera Shooting and Object Interactions. Metrics prefixed with R- are computed within the controlled foreground region. Subj., Bkg., and IQ denote VBench subject consistency, background consistency, and imaging quality, respectively.
Method
R-LPIPS ↓
R-PSNR ↑
Subj. ↑
Bkg. ↑
IQ ↑
GenCine (w/o Fbg )
0.235
16.20
0.889
0.937
0.683
GenCine (w/o Ffg )
0.223
15.26
0.912
0.938
0.679
GenCine ( sside=1 )
0.189
17.17
0.897
0.942
0.686
GenCine ( sside=2 )
0.184
17.16
0.902
0.944
0.691
Table 2 : Guidance-map and side-branch resolution ablations on Camera Shooting examples.
Figure 6 : Control map ablation for GenCine. The target is a fixed camera with the boat following the purple 3D trajectory. Foreground XYZ and identity maps alone (top) show background drift. Background XYZ alone (middle) provides no explicit object motion control. Full guidance (bottom) combines all three maps, better preserving the static background while guiding the boat.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Comparison on an ambiguous motion-control example where the intended edit moves the teapot toward the camera while preserving its orientation in world space. Image-plane trajectories do not uniquely specify the underlying 3D transformation. Wan-Move ( Chu et al., 2025 ) produces geometrically inconsistent motion in this example, while GenCine better preserves the intended orientation and scene geometry using handles authored in the canonical world frame.
Figure 8 : Qualitative comparison across camera and object motion sequences. Each scene is sampled at six aligned timesteps for ATI ( Wang et al., 2025a ) , Wan-Move ( Chu et al., 2025 ) , GenCine, and ground truth. GenCine follows the reference foreground motion while preserving coherent scene context. In the dog sequence, its generated poses more closely match the reference poses than the baseline outputs.
Figure 9 : Examples of the three guidance maps. All maps use the same rasterized image-plane interface. (a) Foreground identity maps overlaid on RGB frames. Each artist-selected 3D handle region is rendered with a fixed identity color through its projected Gaussian footprint. (b) Foreground XYZ maps encode the current 3D positions of handle points as RGB values. (c) Background XYZ maps encode the world positions of valid static points. Both XYZ maps share one per-axis normalization, placing their values in the same coordinate system.
Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The mainstream approach to achieving precise camera control is warping a 3D representation according to the target trajectory. However, such methods fail to fully leverage the 3D priors of video diffusion models (VDMs) and often fall into the Inpainting Trap, resulting in subject inconsistency and degraded generation quality. To address this problem, we propose DepthDirector, a video re-rendering framework with precise camera controllability. By leveraging the depth video from explicit 3D representation as camera-control guidance, our method can faithfully reproduce the dynamic scene of an input video under novel camera trajectories. Specifically, we design a View-Content Dual-Stream Condition mechanism that injects both the source video and the warped depth sequence rendered under the target viewpoint into the pretrained video generation model. This geometric guidance signal enables VDMs to comprehend camera movements and leverage their 3D understanding capabilities, thereby facilitating precise camera control and consistent content generation. Next, we introduce a lightweight LoRA-based video diffusion adapter to train our framework, fully preserving the knowledge priors of VDMs. Additionally, we construct a large-scale multi-camera synchronized dataset named MultiCam-WarpData using Unreal Engine 5, containing 8K videos across 1K dynamic scenes. Extensive experiments show that DepthDirector outperforms existing methods in both camera controllability and visual quality. Our code and dataset will be publicly available.
Dong-Yu Chen, Yixin Guo, Shuojin Yang +2
BNRist, Department of Computer Science and Technology, Tsinghua University Beijing 100084, China
Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approaches still lack precise and unified controllability over both object motion and camera motion within a single generation process. We present UniCaMo, a unified framework that enables simultaneous control of object trajectories and camera viewpoints by directly constructing the input noise of the diffusion model. Specifically, UniCaMo builds a shared 3D-grounded motion-consistent noise space across latent video frames. Sparse 3D point tracks are used to warp the Gaussian noise of the reference frame along desired object trajectories, while a virtual spherical noise representation provides globally consistent noise values for newly revealed scene regions under camera motion. By combining local track-guided noise warping with global sphere-based noise sampling, UniCaMo maintains geometric and temporal consistency under both object movement and viewpoint changes. Because UniCaMo modifies only the input noise, it requires no auxiliary adapters, control branches, or architectural changes to the underlying video diffusion model. With lightweight LoRA fine-tuning on large pretrained video diffusion models, including Wan 2.1 (14B), UniCaMo achieves state-of-the-art results in both video quality and motion controllability on standard controllable video generation benchmarks.
Long Vu, Tan Ngo, Animesh Karnewar +5
Qualcomm AI Research · Trinity College Dublin, Ireland
Video generation has achieved remarkable progress in visual fidelity and controllability, enabling conditioning on text, layout, or motion. Among these, motion control - specifying object dynamics and camera trajectories - is essential for composing complex, cinematic scenes, yet existing interfaces remain limited. We introduce LAMP that leverages large language models (LLMs) as motion planners to translate natural language descriptions into explicit 3D trajectories for dynamic objects and (relatively defined) cameras. LAMP defines a motion domain-specific language (DSL), inspired by cinematography conventions. By harnessing program synthesis capabilities of LLMs, LAMP generates structured motion programs from natural language, which are deterministically mapped to 3D trajectories. We construct a large-scale procedural dataset pairing natural text descriptions with corresponding motion programs and 3D trajectories. Experiments demonstrate LAMP's improved performance in motion controllability and alignment with user intent compared to state-of-the-art alternatives establishing the first framework for generating both object and camera motions directly from natural language specifications. Code, models and data are available on our project page.
Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra +3
Koç University · University College London · Adobe Research +1