Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The mainstream approach to achieving precise camera control is warping a 3D representation according to the target trajectory. However, such methods fail to fully leverage the 3D priors of video diffusion models (VDMs) and often fall into the Inpainting Trap, resulting in subject inconsistency and degraded generation quality. To address this problem, we propose DepthDirector, a video re-rendering framework with precise camera controllability. By leveraging the depth video from explicit 3D representation as camera-control guidance, our method can faithfully reproduce the dynamic scene of an input video under novel camera trajectories. Specifically, we design a View-Content Dual-Stream Condition mechanism that injects both the source video and the warped depth sequence rendered under the target viewpoint into the pretrained video generation model. This geometric guidance signal enables VDMs to comprehend camera movements and leverage their 3D understanding capabilities, thereby facilitating precise camera control and consistent content generation. Next, we introduce a lightweight LoRA-based video diffusion adapter to train our framework, fully preserving the knowledge priors of VDMs. Additionally, we construct a large-scale multi-camera synchronized dataset named MultiCam-WarpData using Unreal Engine 5, containing 8K videos across 1K dynamic scenes. Extensive experiments show that DepthDirector outperforms existing methods in both camera controllability and visual quality. Our code and dataset will be publicly available.
Figures & tables
Figure 1 : Example results synthesized by DepthDirector. DepthDirector re-shoots the source video with novel camera trajectories.
Figure 2 : Limitations of warping-based methods . Reprojected pixels exhibit noisy artifacts due to inaccurate 3D geometry. It leads to unrecoverable distortion to the identity and details of the subject, especially on human faces(See Red Box). Reflections on the table are baked as textures(Yellow Box). These inpainting artifacts are attributed to the unawareness of camera movements and video content.
Figure 3 : Architecture overview of DepthDirector . We render depth video and occlusion mask video under target camera viewpoint from an explicit 3D mesh, injecting it into the noise latent by projection and addition. Source video is frame-wise concatenated alongside to provide content reference. Then a LoRA-based video diffusion adapter is trained to generate video following novel camera trajectories.
Figure 4 : Illustration of our dataset: MultiCamWarp Dataset.
Figure 5 : Qualitative comparison with state-of-the-art methods. Ours DepthDirector achieves both stable camera controllability and content preservation. GEN3C [ 38 ] and EX-4D [ 20 ] generate distorted artifact due to inaccuracy of depth estimation (see zoom-out details within the yellow border). RCM(ReCamMaster) [ 3 ] not only loses consistency at large rotation angle, but also fails to achieve stable camera control. It fails to return back to the original viewpoint in the “orbit” camera trajectory (see red dot on the right).
Figure 6 : Qualitative comparison with warping-based methods . Our method preserves facial identity under camera view changes, while other methods produce noticeable distortions.
Method
Camera Accuracy
Identity Preservation
VBench
RE ↓
TE ↓
CamMC ↓
RS ↑
IFS ↑
Cs.S. ↑
Cs.B. ↑
Img.Q ↑
M.S. ↑
TC [ 61 ]
1.464
0.133
2.084
0.567
0.916
93.85
94.75
71.59
98.37
GEN3C [ 38 ]
1.397
0.076
1.982
0.618
0.943
94.50
93.80
72.04
99.21
EX4D [ 20 ]
1.186
0.101
1.687
0.653
0.925
94.18
94.58
70.65
98.12
Ours(DC)
1.382
0.096
1.962
0.669
0.959
94.62
94.76
72.88
99.28
RCM [ 3 ]
3.576
0.180
5.063
0.576
0.932
92.78
91.15
67.27
99.17
Table 1 : Quantitative comparison with state-of-the-art methods. The best numbers are bolded and the second ones are unlerlined.
Figure 7 : Qualitative comparison with implicit-controlled methods . Our method generates consistent background, while baseline methods fail to preserve the environment details from input views.
Figure 8 : Qualitative comparison on Extreme Novel View. DepthDirector can achieve high-quality video generation results with stable camera controllability even with camera rotations of up to 90 degrees.
Method
Mild ( 0∘±30∘ )
Large ( 0∘±60∘ )
Extreme ( 0∘±90∘ )
Cs.S. ↑
Cs.B. ↑
Img.Q. ↑
Cs.S. ↑
Cs.B. ↑
Img.Q. ↑
Cs.S. ↑
Cs.B. ↑
Img.Q. ↑
EX-4D [ 20 ]
94.37
94.58
70.77
90.09
91.85
69.43
86.03
89.02
67.30
GEN3C [ 38 ]
94.43
94.05
72.98
90.97
92.12
70.75
87.73
90.07
67.64
RCM [ 3 ]
93.81
91.64
68.83
89.79
89.54
64.79
83.54
86.31
58.42
CamClone [ 36 ]
95.35
93.63
70.46
92.63
91.83
69.99
90.04
90.46
69.32
Ours(Full)
94.70
95.22
72.92
91.91
92.93
72.66
89.92
90.91
71.83
Table 2 : Quantitative comparison on Extreme Novel View.
Figure 9 : Ablation Study on Model Design. This demonstrates the importance of decomposed condition injection by depth condition and View-Content Dual-Stream mechanism.
Method
Identity
VBench
RS ↑
IFS ↑
Cs.S. ↑
Cs.B. ↑
Model (a)
0.5954
0.9606
94.80
94.57
Model (b)
0.6463
0.9621
95.20
94.43
Model (c)
0.5804
0.9628
95.17
94.66
Full Model
0.6887
0.9661
95.29
94.66
Table 3 : Ablation study on Model Design.
Figure 10 : Ablation Study on Model Design. Model without injecting source video cannot recover the dynamic expression change of source video.
Figure 11 : Attention visualization. Green box marks the query noise patch; the heatmap shows how source video patches attended by it. We normalize weight by a factor(0.04), red means higher. DepthDirector maintains strong and precise attentions to corresponding patch from source video, producing consistent outputs, while the RGB variant produces weak and rough attentions, leading to distorted structures.
Figure B1 : Reference Video for CamCloneMaster
Method
R-Err ↓
T-Err ↓
Cs.Subj. ↑
Cs.Bg. ↑
ReCamMaster dataset
1.88
0.12
94.96
94.97
Full (MultiCamWarp Dataset)
1.15
0.08
95.03
94.90
Table A1 : Ablation Study on Dataset
Method
Resolution
Frames
Denoising Steps
Inference Time
Base Model
Inference Time(Base)
EX-4D
480×832
49
25
305s
Wan2.1-I2V-14B-480P
250s
TrajectoryCrafter
384×672
49
50
160s
CogVideoX-Fun-5B
150s
GEN3C
704×1280
121
35
840s
Cosmos-7B-Video2World
780s
ReCamMaster
480×832
81
50
518s
Wan2.1-T2V-1.3B
181s
CamCloneMaster
480×832
81
50
1126s
Wan2.1-T2V-1.3B
181s
Ours(DepthDirector)
704×1280
81
50
240s
Wan2.2-TI2V-5B
180s
Table A2 : Efficiency Comparison
Method
RE
TE
CamMC
RS
IFS
Cs.S.
Cs.B.
Img.Q
M.S.
DaS [ 15 ]
3.806
0.294
5.095
0.564
0.916
92.37
92.58
67.84
98.81
Table A3 : Quantitative Results of DaS [ 15 ]
Figure B2 : Different Choises of Video Depth Estimation. We visualize the input frame normal map derived from estimated depth map, warped depth map, warped video based on estimated depth map and the generated results of our methods and warping-based methods. Due to the inherent inaccuracy of monocular depth estimation, the warp results from these SOTA video estimation methods contain distortion and artifacts without exception.
Figure B3 : Robustness. We add Gaussian noise (std =10% of mean depth) to the estimated depth map before warping the geometry into novel camera trajectories. The generated video still remains temporally consistent and follows the target camera viewpoints.
Method
Mild ( 0∘±30∘ )
Large ( 0∘±60∘ )
Extreme ( 0∘±90∘ )
RotErr ↑
TErr ↑
CamMC ↑
RotErr ↑
TErr ↑
CamMC ↑
RotErr ↑
TErr ↑
CamMC ↑
RCM [ 3 ]
4.809
0.104
6.799
6.985
0.105
9.857
8.493
0.154
11.991
CamClone [ 36 ]
5.543
0.174
7.834
8.556
0.094
12.077
11.45
0.093
16.155
Ours(Full)
1.281
0.032
1.813
2.499
0.043
3.534
4.031
0.035
5.695
Table A4 : Stress Test On Extreme Viewpoint Change
Figure B4 : More Results of our method.
Figure B5 : More Comparison of Extreme Novel View.
Figure B6 : More Comparison of Extreme Novel View.
Figure B7 : More Comparison of Extreme Novel View.
Figure B8 : More Comparison of Extreme Novel View.
Video is a rich and scalable source of 3D/4D visual observations, and camera control is a key capability for video generation models to produce geometrically meaningful content. Existing approaches typically learn a mapping from camera motion to video using additional camera modules and paired data. However, such datasets are often limited in scale, diversity, and scene dynamics, which can bias the model toward a narrow output distribution and compromise the strong prior learned by the base model. These limitations motivate a different perspective on camera control. In this paper, we show that camera control need not be modeled as an implicit mapping problem, but can instead be treated as a form of geometric guidance that induces displacements across frames. Specifically, we reformulate camera control into a set of displacement fields and apply them via differentiable resampling of latent features during denoising. Our simple approach achieves effective camera control with minimal degradation across diverse quality metrics compared to fine-tuned baselines. Since our method is applicable to most video diffusion models without training, it can also serve as a probe to study the camera control capabilities of base models. Using this probe, we identify universal biases shared by representative video models, as well as disparities in their responses to camera control. Finally, we benchmark their performance in multi-view generation, offering insights into their potential for 3D/4D tasks.
Chen Hou, Christian Rupprecht
Visual Geometry Group · Visual Geometry Group University of Oxford
Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approaches still lack precise and unified controllability over both object motion and camera motion within a single generation process. We present UniCaMo, a unified framework that enables simultaneous control of object trajectories and camera viewpoints by directly constructing the input noise of the diffusion model. Specifically, UniCaMo builds a shared 3D-grounded motion-consistent noise space across latent video frames. Sparse 3D point tracks are used to warp the Gaussian noise of the reference frame along desired object trajectories, while a virtual spherical noise representation provides globally consistent noise values for newly revealed scene regions under camera motion. By combining local track-guided noise warping with global sphere-based noise sampling, UniCaMo maintains geometric and temporal consistency under both object movement and viewpoint changes. Because UniCaMo modifies only the input noise, it requires no auxiliary adapters, control branches, or architectural changes to the underlying video diffusion model. With lightweight LoRA fine-tuning on large pretrained video diffusion models, including Wan 2.1 (14B), UniCaMo achieves state-of-the-art results in both video quality and motion controllability on standard controllable video generation benchmarks.
Long Vu, Tan Ngo, Animesh Karnewar +5
Qualcomm AI Research · Trinity College Dublin, Ireland
Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme scarcity of synchronized, multi-view real-world video data. Consequently, the prevailing paradigm often exhibits limited generalization when processing out-of-distribution real-world videos, with models struggling to accurately adhere to physical scales and camera trajectories. To bridge this gap, we propose Geo-Align, the first Reinforcement Learning framework specifically designed for camera-controlled video re-rendering. Built upon a pretrained model, we optimize the model through a scale-aware perceptual reward mechanism. Specifically, we introduce a metric 3D estimator to extract precise camera trajectories from generated videos, explicitly penalizing deviations in rotation and translation. Furthermore, we meticulously designed a data pipeline strategy based on real-world conditioning videos and target camera trajectories derived from synthetic data, eliminating the reliance on paired data. Extensive experiments demonstrate that Geo-Align consistently outperforms existing supervised learning baselines in both precise camera controllability and visual fidelity, indicating the effectiveness of our method.