Controlling the camera relative to a moving subject in an existing video is challenging: behaviors such as maintaining a frontal view require the camera to adapt to the subject's changing position and orientation, making the desired trajectory difficult to specify in advance. Existing camera-controlled video-to-video methods typically rely on explicit trajectories or reference motions, which do not directly express these dynamic camera--subject relationships. We introduce semantic camera motion control, a novel video-to-video task in which a reference video and a target motion label specify the desired subject-relative camera behavior without an explicit target trajectory. Our method, SemCam, learns to realize this behavior while preserving source content. It combines shared-basis low-rank adaptation with motion-conditioned modulation, while a background-consistency loss encourages fidelity in regions visible in both reference and target videos. We construct 661 paired videos covering eight semantic camera behaviors and evaluate on a separate 109-scene benchmark using subject-relative motion metrics, appearance measures, and a user study. SemCam achieves a semantic-motion success rate of 68.6%, compared with 45.3% for Vista4D, the strongest evaluated baseline, while maintaining comparable subject identity preservation.
Figures & tables
Figure 2: SemCam method overview. Reference-video tokens are concatenated with noisy target-video tokens zt and processed by the DiT. The motion label m selects a motion-specific LoRA factor Ak , paired with a shared basis B , and a learned embedding em that predicts modulation residuals (Δγm,Δβm,Δgm) . These residuals are added to the timestep-dependent modulation parameters. Fire icons indicate trainable components; the pretrained backbone remains frozen.
Motion
Subject appearance
VLM judge
Method
Sem. motion ↑
Success ↑
DINO ↑
ID-Sim ↓
Sem. motion ↑
Identity ↑
TrajectoryCrafter
0.238
0.223
0.426
0.506
0.377
0.520
ReCamMaster
0.139
0.138
0.531
0.407
0.281
0.725
CamCloneMaster
0.180
0.172
0.591
0.329
0.299
0.732
Vista4D
0.429
0.453
0.592
0.307
0.476
0.750
Ours
0.642
0.686
0.594
0.293
0.621
0.750
Table 1: Comparison with camera-control baselines on 109 scenes and eight semantic camera target motions (872 clips per method). Scores are averaged within each motion type and then across motions. Sem. motion denotes the execution score M ; Success is the fraction of clips with M>0.5 . DINO and ID-Sim measure subject appearance relative to the input, while the VLM judge evaluates motion execution and identity preservation accounting for viewpoint changes.
Figure 3: Qualitative comparison with camera-control baselines. Four examples compare SemCam (highlighted in blue) with Vista4D [ LLS∗26 ] , TrajectoryCrafter [ YHXS25 ] , ReCamMaster [ BXF∗25 ] , and CamCloneMaster [ LSB∗25 ] . Each example is labeled by the requested camera behavior; rows show the input and method outputs, and columns show four frames sampled over time.
Figure 4: Semantic camera control across diverse scenes. Each cell shows three frames from the input video (top, blue) and SemCam ’s output (bottom) for the labeled camera behavior. Outputs are generated independently from each input video and motion label. Examples span people, animals, vehicles, and stylized subjects; all clips are included in the supplementary videos.
Figure 5: User study comparing our method against Vista4D [ LLS∗26 ] (top) and CamCloneMaster [ LSB∗25 ] (bottom), conducted with 55 participants. Each horizontal bar shows the fraction of responses preferring our method ( Ours), rating both equally ( Both), or preferring the baseline ( Baseline) for camera-control alignment (top) and visual quality (bottom).
Figure 6: Ablation study. Two semantic camera motions (columns; two frames per clip) for the full model and each ablation of Tab. 3 . Our full model both executes the requested camera motion control and preserves the scene.
Method
Follow
Top
Front
Behind
Back → front
Front → back
Zoom face
Zoom figure
TrajectoryCrafter
0.477
0.179
0.284
0.138
0.099
0.071
0.156
0.377
ReCamMaster
0.495
0.075
0.303
0.037
0.039
0.010
0.093
0.056
CamCloneMaster
0.477
0.064
0.220
0.028
0.196
0.058
0.220
0.111
Vista4D
0.771
0.587
0.303
0.156
0.472
0.288
0.519
0.528
Ours
0.743
0.761
0.817
0.743
0.575
0.500
0.761
0.587
Table 2: Success rate per semantic motion ( M>0.5 ), 109 scenes per motion. Best per column in bold .
Motion
Subject appearance
VLM judge
Variant
Sem. motion ↑
Success ↑
DINO ↑
ID-Sim ↓
Sem. motion ↑
Identity ↑
Full (Ours)
0.654
0.706
0.636
0.283
0.610
0.737
− background loss
0.611
0.634
0.638
0.284
0.602
0.720
− motion modulation
0.530
0.554
0.662
0.259
0.551
0.768
Shared LoRA ( AB )
0.306
0.277
0.704
0.225
0.436
0.809
Per-motion LoRA ( AiBi )
0.455
0.447
0.642
0.276
0.540
0.742
Table 3: Ablation of design choices on a subset of 33 scenes. Metrics and aggregation follow Tab. 1 . Appearance scores should be interpreted alongside motion success, as outputs with little camera movement can retain high reference similarity. Best values in each column are shown in bold . Layer-selection and rank ablations appear in Tab. A4 .
Figure 7: Motion fusion. Weighted combinations of motion-specific LoRA factors produce composite camera behaviors without additional training. Left: Move up . Middle: Side rotation (top) and follow (bottom). Right: The corresponding combinations, showing upward camera movement combined with rotation (top) or subject-following (bottom).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Where the background-consistency loss looks. One row per semantic motion: reference and ground-truth target frames, the loss region M=covis∧¬subject , and the per-patch cosine agreement between the proxy features g(z^0) and DINO (x⋆) , shown only on M (green = match, red = mismatch). The mask excludes the subject and the regions the reference never saw, so the loss is applied only where the background should persist.
#
Motion type
Pairs
1
Follow
97
2
Follow front view
103
3
Follow top view
84
4
Follow behind
86
5
Rotate back → front
108
6
Rotate front → back
100
Appendix
Table A1: Training pairs for the eight semantic camera motions.
Motion
Components ( M=min )
What they check
follow
Travel, Dist. kept
camera moves with the subject at a constant distance
follow top
Travel, Dist. kept, Elevation
as follow, from above the subject
follow front
Travel, Dist. kept, View (front)
as follow, facing the subject’s front throughout
follow behind
Travel, Dist. kept, View (back)
as follow, behind the subject throughout
rotate back → front
View (back → front), Dist. kept
orbits from behind to the front at a constant radius
rotate front → back
View (front → back), Dist. kept
orbits from the front to behind at a constant radius
Appendix
Table A2: Semantic-motion score per motion type : M is the minimum of the listed components.
Method
Travel
Dist. kept
Elevation
View
Zoom
Face at end
Framing
TrajectoryCrafter
0.611
0.624
0.290
0.256
0.384
0.530
0.681
ReCamMaster
0.599
0.481
0.166
0.180
0.132
0.233
0.115
CamCloneMaster
0.647
0.546
0.170
0.232
0.300
0.582
0.493
Vista4D
0.703
0.756
0.716
0.356
0.595
0.719
0.909
Ours
0.843
0.800
0.824
0.726
0.683
0.823
0.826
Appendix
Table A3: Breakdown of the semantic-motion score on the 109 scenes; each component is averaged over the motions that use it (Tab. A2 ).
Figure A2: Our dataset. Examples of input–target video pairs used for training. Each pair consists of a reference video and a corresponding target video exhibiting a specific camera transformation.
Figure A3: Extended qualitative results. Additional SemCam generations (each cell: input video top row, SemCam output bottom row; three frames over time) across further scenes and camera motions. As in Fig. 4 , each clip is an independent video-to-video generation from the input plus a single motion token.
Figure A4: Additional qualitative comparisons with camera-control baselines. SemCam (highlighted in blue) against Vista4D, TrajectoryCrafter, ReCamMaster, and CamCloneMaster on further scenes, one block per camera behavior; rows show the input and method outputs, columns four frames sampled over time.
Figure A5: Additional qualitative comparisons with camera-control baselines (continued). SemCam (highlighted in blue) against Vista4D, TrajectoryCrafter, ReCamMaster, and CamCloneMaster on further scenes, one block per camera behavior; rows show the input and method outputs, columns four frames sampled over time.
Figure A6: Screenshot of the Google Forms interface shown to participants in the user study.
Figure A7: Example of a user study question showing intended camera motion, reference video (Ref), generated video options (A and B), and the corresponding evaluation criteria (Camera-Control Alignment and Visual Alignment).
Motion
Subject appearance
VLM judge
Variant
Sem. motion ↑
Success ↑
DINO ↑
ID-Sim ↓
Sem. motion ↑
Identity ↑
Design components
Full (Ours)
0.654
0.706
0.636
0.283
0.610
0.737
− background loss
0.611
0.634
0.638
0.284
0.602
0.720
− motion modulation
0.530
0.554
0.662
0.259
0.551
0.768
Shared LoRA ( AB )
0.306
0.277
0.704
0.225
0.436
0.809
Appendix
Table A4: Full ablation on the 33-scene subset. Best per column over all variants in bold .
Camera motion control is essential for directing viewpoint changes in generative systems. However, existing methods typically condition the generation process on a single specific modality, such as explicit pose trajectories or reference videos, limiting their ability to support heterogeneous user inputs. To address this limitation, we present TriMotion, a modality-agnostic framework for camera-controlled video generation that maps video, pose, and text inputs, describing the same camera trajectory into a shared motion embedding space. Learning such a space requires synchronized supervision across modalities. Therefore, we build the Motion Triplet Dataset by extending a Multi-Cam Video Dataset with geometry-grounded motion descriptions derived from camera extrinsics. We further introduce a latent motion consistency objective that leverages the motion embedding space to encourage the generated video to follow the target camera trajectory directly in latent space, avoiding the cost of pixel-space decoding. Extensive experiments show that TriMotion generates high-quality videos that accurately follow the target camera trajectories across all three modalities. Beyond standard generation, the shared motion embedding space also enables flexible applications such as sequential motion composition and cross-modal motion interpolation.
Seunghyun Shin, Jifei Song, Wooseok Jeon +2
GIST · Huawei Noah’s Ark Lab · Yonsei University +1
For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.
Omar El Khalifi, Thomas Rossi, Oscar Fossey +6
Kinetix, France · University of Oxford, United Kingdom · MBZUAI, United Arab Emirates
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typically rely on visual motion sequences, such as skeleton maps, pose maps, or rendered body representations, for motion control, while using camera embeddings for camera control. Such heterogeneous control interfaces force video generation models to reconcile pixel-aligned visual cues with non-visual geometric embeddings, making motion-camera attribution difficult and sensitive to camera estimation errors. We propose \textbf{UniMoCa}, a representation-driven framework that unifies motion and camera controls in visual space. At the core of UniMoCa is \textbf{Motion-Camera Visual Proxy} (\textbf{MCVP}), a mutually-sharable novel representation that converts 3D human motion and camera trajectories extracted from driving videos into an identity-neutral visual proxy. MCVP renders temporally aligned human geometry under the recovered camera trajectory and augments it with explicit camera trajectory markers, replacing heterogeneous visual-parametric controls with distinguishable visual cues. As both control factors are represented in the same visual space, they become mutually compatible rather than heterogeneous, enabling consistent joint reasoning and editing during video generation. We further curate a \textbf{MCVP-Video} dataset covering complex actions, multi-person interactions, and diverse camera trajectories. Experiments based on the Wan2.2 I2V show that UniMoCa achieves substantial gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity. More details are shown in our Project page: https://tanliming-daniel.github.io/UniMoCa/.
Liming Tan, Ye Chen, Hao Zhang +3
Shanghai Jiao Tong University · Monash University · USC-SJTU Institute of Cultural and Creative Industry