Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.
Figures & tables
Figure 1: Camera and object control from a single image. Top: the image is lifted into a background point cloud and complete rigid 3D geometry for each object, in which the user draws the object and camera trajectories and, optionally, inserts a new object from a reference image. Bottom: 4Director generates videos that follow both trajectories, with plausible object motion, consistent appearance and illumination, and the background revealed by the camera filled in.
Figure 2: Control representations. Top: given an input image and an object trajectory (a 180∘ turn in place), we compare four control signals: a 2D bounding box, a 3D Gaussian blob, 3D point trajectories, and our rigid 3D geometry. Bottom: capabilities of the compared representations ( ✓ supported, ✓ p partial, ✗ not supported, – not stated in the paper).
Figure 3: Overview of 4Director. Left: each object is reconstructed once as a complete canonical mesh and moved by one rigid transformation per frame; the meshes, a background point cloud and the camera share one coordinate frame, in which the user prescribes the trajectory of each object and the camera, and the scene is rendered to a depth video (Sec. 3.1 ). Right: the Motion Adapter , a trainable branch, injects this rigid rendering into a pretrained video generator, which supplies appearance, illumination and non-rigid dynamics (Sec. 3.2 ); it is trained on RealCOD-Rigid (Sec. 4 ).
Figure 4: RealCOD-Rigid annotation pipeline. From each clip we estimate object masks, depth, and the camera trajectory, which together give a dynamic 4D point cloud of the scene ( 4D Point Cloud – Dynamic ); we reconstruct a canonical mesh for every masked object from the first frame ( First Frame to 3D ) and track 3D points on it as a rigid body ( Rigid Body Tracking ), which yields one rigid transformation per frame and aligns the canonical mesh to every frame ( 4D Point Cloud – Rigid 3D Geometry ), giving the rigid 3D scene. Rendering it to a depth video yields the control; the original clip is the training target, with its text description as the prompt.
Table 1: Quantitative comparison on joint camera and object control. (a) Visual quality and text alignment (FID, FVD, CLIP-SIM), camera control (RotErr, TransErr) and object control (recognition rate, IG-IoU). Best bold , second best underlined . (b) Identity-Gated IoU on one frame: a vision–language judge decides whether the generated object is still the same object; a valid frame contributes its mask IoU and an invalid frame zero (Eq. 6 ).
VBench-I2V
User study (1–5)
Method
Subject
Background
Motion
Dynamic
Aesthetic
Imaging
I2V
I2V
Quality
Control
Consistency
consistency
consistency
smoothness
degree
quality
quality
subject
background
↑
↑
↑
MotionCtrl
70.8
84.3
92.9
94.0
43.8
45.7
81.6
84.2
1.49
1.44
1.79
PaC
90.0
91.6
95.1
93.0
55.0
61.6
92.4
92.5
1.98
2.01
2.41
VerseCrafter
88.9
91.9
98.7
91.0
56.7
64.2
96.0
96.8
3.69
2.39
3.88
SymphoMotion
89.5
91.4
98.5
91.0
54.8
59.1
95.4
96.7
3.35
2.33
3.69
Table 2: VBench-I2V and user study. VBench-I2V scores on the eight applicable dimensions (percent) and mean user ratings of visual quality, control accuracy and consistency (1 to 5, Sec. 5.3 ). 4Director leads on every dimension; VBench margins over the next best stay within 2 points, whereas the user study separates them clearly. Best bold , second best underlined .
Figure 5: Qualitative comparison on joint camera and object control. For each clip, the top row shows the input image and the prescribed trajectories, and the rows below show temporally aligned frames from the four baselines, 4Director and the source video. Arrows mark the object heading, green in the source and red in the generated videos. Only 4Director completes the turn of the source video; the baselines keep the original heading, turn only partway, blur or lose the subject.
Figure 6: Objects leaving and re-entering the view. Top left: the partial point cloud lifted from the input image; the camel walks out of the camera’s view and comes back (blue arrows) while the camera stays fixed. For each method, the upper row is the control it receives and the lower row the video generated from it. Only 4Director shows the empty frame while the camel is away and then brings the same camel back.
Figure 7: Ablation on the control representation. Top: the source video and three alternative controls, each shown as the control video above the video generated from it; arrows mark the motorcycle heading, and only the rigid 3D geometry follows the prescribed turn. Bottom: what each control encodes and its scores; the three arms are trained with the recipe of 4Director, and the last row is 4Director as in Table 1(a) . Best bold , second best underlined .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
4Director (full)
4Director-LoRA
Trainable parameters
entire adapter, 3.049 B
rank-16 LoRA on the context blocks’ attention and feed-forward projections, ≈15.3 M
Initialization
official VACE + NoisyTune-style noise
official VACE
GPUs / global batch
3×8 / 24
3×8 / 24
Epochs
3
3
Peak learning rate
5×10−5
5×10−5
Warmup steps
25
25
Appendix
Table 3: Training hyper-parameters of the two variants.
FID
FVD
CLIP
Rot. ( ∘ )
Trans.
Rec.
IG-IoU
Subj.
Bkg.
Motion
Dyn.
Aesth.
Imag.
I2V
I2V
Method
↓
↓
↑
↓
↓
↑
↑
cons.
cons.
smooth.
degree
quality
quality
subj.
bkg.
4Director-LoRA
44.4
393.2
31.4
4.02
0.131
94.1
57.3
90.0
91.5
98.7
92.0
56.8
63.2
96.2
96.9
4Director
44.1
370.4
31.8
3.65
0.122
94.9
60.4
90.3
92.0
98.8
96.0
57.0
64.6
96.3
97.0
Appendix
Table 4: LoRA variant against the full adapter. The metrics of Table 1(a) (left) and the VBench-I2V dimensions of Table 2 (right).
Figure 8: Occlusion. A plane is directed to fly under the bridge and behind its right tower. Rows 1–2: the control of SymphoMotion, a 2D box projected from the 3D trajectory onto the rendered point cloud (red), and the video it generates. Rows 3–4: the depth control rendered from the rigid 3D scene and the video 4Director generates from it. Only with the depth control does the tower hide the plane.
Figure 9: Failure case. A breakdancer in a handstand is moved to the right as one rigid body while the camera moves. Top: the rigid 3D scene rendered to RGB, for visualization only, and to the depth video that conditions the adapter; pixels without geometry are black in the RGB render and uniform gray in the depth video. Bottom: the generated dancer keeps the pose of the first frame throughout, following the rigid silhouette, instead of dancing.
Figure 10: Authoring interface. The three stages of building and directing a rigid 3D scene from a single image; in (c) the trajectories are set by keyframes on a timeline.
Figure 11: One page of the user study. Left: the case and the five anonymized results. Right: the three rating questions.
Figure 12: Authored camera and object trajectories (1/2). Each panel shows the input image and the authored trajectories on the left, the depth control rendered from the rigid 3D scene in the upper row, and the generated video below it. None of these controls comes from a source video.
Figure 13: Authored camera and object trajectories (2/2). As in Fig. 12 . The generator supplies the non-rigid dynamics, illumination and background that the rigid control leaves unspecified.
Figure 14: Camera and object motion. In the first two clips, the camera follows a trajectory (the coloured frusta) while the object is moved by its own rigid trajectory. In the third, the camera follows an arc around the subject and the object keeps its place, so the non-rigid dynamics come from the generator alone.
Figure 15: Object rotation with a static camera. The object is turned by 180∘ in place, with the camera held at the single frustum on the left. The baselines keep the original heading or lose the subject, whereas 4Director completes the turn and shows the side that turns into view.
Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The mainstream approach to achieving precise camera control is warping a 3D representation according to the target trajectory. However, such methods fail to fully leverage the 3D priors of video diffusion models (VDMs) and often fall into the Inpainting Trap, resulting in subject inconsistency and degraded generation quality. To address this problem, we propose DepthDirector, a video re-rendering framework with precise camera controllability. By leveraging the depth video from explicit 3D representation as camera-control guidance, our method can faithfully reproduce the dynamic scene of an input video under novel camera trajectories. Specifically, we design a View-Content Dual-Stream Condition mechanism that injects both the source video and the warped depth sequence rendered under the target viewpoint into the pretrained video generation model. This geometric guidance signal enables VDMs to comprehend camera movements and leverage their 3D understanding capabilities, thereby facilitating precise camera control and consistent content generation. Next, we introduce a lightweight LoRA-based video diffusion adapter to train our framework, fully preserving the knowledge priors of VDMs. Additionally, we construct a large-scale multi-camera synchronized dataset named MultiCam-WarpData using Unreal Engine 5, containing 8K videos across 1K dynamic scenes. Extensive experiments show that DepthDirector outperforms existing methods in both camera controllability and visual quality. Our code and dataset will be publicly available.
Dong-Yu Chen, Yixin Guo, Shuojin Yang +2
BNRist, Department of Computer Science and Technology, Tsinghua University Beijing 100084, China
Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.
We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results demonstrate that our method supports the synthesis of complex and extended events with unprecedented controllability and persistent dynamic object memory. Project Page: https://worlddirector.github.io/