Controlling the camera relative to a moving subject in an existing video is challenging: behaviors such as maintaining a frontal view require the camera to adapt to the subject's changing position and orientation, making the desired trajectory difficult to specify in advance. Existing camera-controlled video-to-video methods typically rely on explicit trajectories or reference motions, which do not directly express these dynamic camera--subject relationships. We introduce semantic camera motion control, a novel video-to-video task in which a reference video and a target motion label specify the desired subject-relative camera behavior without an explicit target trajectory. Our method, SemCam, learns to realize this behavior while preserving source content. It combines shared-basis low-rank adaptation with motion-conditioned modulation, while a background-consistency loss encourages fidelity in regions visible in both reference and target videos. We construct 661 paired videos covering eight semantic camera behaviors and evaluate on a separate 109-scene benchmark using subject-relative motion metrics, appearance measures, and a user study. SemCam achieves a semantic-motion success rate of 68.6%, compared with 45.3% for Vista4D, the strongest evaluated baseline, while maintaining comparable subject identity preservation.
Figures & tables
Figure 2: SemCam method overview. Reference-video tokens are concatenated with noisy target-video tokens zt and processed by the DiT. The motion label m selects a motion-specific LoRA factor Ak , paired with a shared basis B , and a learned embedding em that predicts modulation residuals (Δγm,Δβm,Δgm) . These residuals are added to the timestep-dependent modulation parameters. Fire icons indicate trainable components; the pretrained backbone remains frozen.
Motion
Subject appearance
VLM judge
Method
Sem. motion ↑
Success ↑
DINO ↑
ID-Sim ↓
Sem. motion ↑
Identity ↑
TrajectoryCrafter
0.238
0.223
0.426
0.506
0.377
0.520
ReCamMaster
0.139
0.138
0.531
0.407
0.281
0.725
CamCloneMaster
0.180
0.172
0.591
0.329
0.299
0.732
Vista4D
0.429
0.453
0.592
0.307
0.476
0.750
Ours
0.642
0.686
0.594
0.293
0.621
0.750
Table 1: Comparison with camera-control baselines on 109 scenes and eight semantic camera target motions (872 clips per method). Scores are averaged within each motion type and then across motions. Sem. motion denotes the execution score M ; Success is the fraction of clips with M>0.5 . DINO and ID-Sim measure subject appearance relative to the input, while the VLM judge evaluates motion execution and identity preservation accounting for viewpoint changes.
Figure 3: Qualitative comparison with camera-control baselines. Four examples compare SemCam (highlighted in blue) with Vista4D [ LLS∗26 ] , TrajectoryCrafter [ YHXS25 ] , ReCamMaster [ BXF∗25 ] , and CamCloneMaster [ LSB∗25 ] . Each example is labeled by the requested camera behavior; rows show the input and method outputs, and columns show four frames sampled over time.
Figure 4: Semantic camera control across diverse scenes. Each cell shows three frames from the input video (top, blue) and SemCam ’s output (bottom) for the labeled camera behavior. Outputs are generated independently from each input video and motion label. Examples span people, animals, vehicles, and stylized subjects; all clips are included in the supplementary videos.
Figure 5: User study comparing our method against Vista4D [ LLS∗26 ] (top) and CamCloneMaster [ LSB∗25 ] (bottom), conducted with 55 participants. Each horizontal bar shows the fraction of responses preferring our method ( Ours), rating both equally ( Both), or preferring the baseline ( Baseline) for camera-control alignment (top) and visual quality (bottom).
Figure 6: Ablation study. Two semantic camera motions (columns; two frames per clip) for the full model and each ablation of Tab. 3 . Our full model both executes the requested camera motion control and preserves the scene.
Method
Follow
Top
Front
Behind
Back → front
Front → back
Zoom face
Zoom figure
TrajectoryCrafter
0.477
0.179
0.284
0.138
0.099
0.071
0.156
0.377
ReCamMaster
0.495
0.075
0.303
0.037
0.039
0.010
0.093
0.056
CamCloneMaster
0.477
0.064
0.220
0.028
0.196
0.058
0.220
0.111
Vista4D
0.771
0.587
0.303
0.156
0.472
0.288
0.519
0.528
Ours
0.743
0.761
0.817
0.743
0.575
0.500
0.761
0.587
Table 2: Success rate per semantic motion ( M>0.5 ), 109 scenes per motion. Best per column in bold .
Motion
Subject appearance
VLM judge
Variant
Sem. motion ↑
Success ↑
DINO ↑
ID-Sim ↓
Sem. motion ↑
Identity ↑
Full (Ours)
0.654
0.706
0.636
0.283
0.610
0.737
− background loss
0.611
0.634
0.638
0.284
0.602
0.720
− motion modulation
0.530
0.554
0.662
0.259
0.551
0.768
Shared LoRA ( AB )
0.306
0.277
0.704
0.225
0.436
0.809
Per-motion LoRA ( AiBi )
0.455
0.447
0.642
0.276
0.540
0.742
Table 3: Ablation of design choices on a subset of 33 scenes. Metrics and aggregation follow Tab. 1 . Appearance scores should be interpreted alongside motion success, as outputs with little camera movement can retain high reference similarity. Best values in each column are shown in bold . Layer-selection and rank ablations appear in Tab. A4 .
Figure 7: Motion fusion. Weighted combinations of motion-specific LoRA factors produce composite camera behaviors without additional training. Left: Move up . Middle: Side rotation (top) and follow (bottom). Right: The corresponding combinations, showing upward camera movement combined with rotation (top) or subject-following (bottom).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Where the background-consistency loss looks. One row per semantic motion: reference and ground-truth target frames, the loss region M=covis∧¬subject , and the per-patch cosine agreement between the proxy features g(z^0) and DINO (x⋆) , shown only on M (green = match, red = mismatch). The mask excludes the subject and the regions the reference never saw, so the loss is applied only where the background should persist.
#
Motion type
Pairs
1
Follow
97
2
Follow front view
103
3
Follow top view
84
4
Follow behind
86
5
Rotate back → front
108
6
Rotate front → back
100
Appendix
Table A1: Training pairs for the eight semantic camera motions.
Motion
Components ( M=min )
What they check
follow
Travel, Dist. kept
camera moves with the subject at a constant distance
follow top
Travel, Dist. kept, Elevation
as follow, from above the subject
follow front
Travel, Dist. kept, View (front)
as follow, facing the subject’s front throughout
follow behind
Travel, Dist. kept, View (back)
as follow, behind the subject throughout
rotate back → front
View (back → front), Dist. kept
orbits from behind to the front at a constant radius
rotate front → back
View (front → back), Dist. kept
orbits from the front to behind at a constant radius
Appendix
Table A2: Semantic-motion score per motion type : M is the minimum of the listed components.
Method
Travel
Dist. kept
Elevation
View
Zoom
Face at end
Framing
TrajectoryCrafter
0.611
0.624
0.290
0.256
0.384
0.530
0.681
ReCamMaster
0.599
0.481
0.166
0.180
0.132
0.233
0.115
CamCloneMaster
0.647
0.546
0.170
0.232
0.300
0.582
0.493
Vista4D
0.703
0.756
0.716
0.356
0.595
0.719
0.909
Ours
0.843
0.800
0.824
0.726
0.683
0.823
0.826
Appendix
Table A3: Breakdown of the semantic-motion score on the 109 scenes; each component is averaged over the motions that use it (Tab. A2 ).
Figure A2: Our dataset. Examples of input–target video pairs used for training. Each pair consists of a reference video and a corresponding target video exhibiting a specific camera transformation.
Figure A3: Extended qualitative results. Additional SemCam generations (each cell: input video top row, SemCam output bottom row; three frames over time) across further scenes and camera motions. As in Fig. 4 , each clip is an independent video-to-video generation from the input plus a single motion token.
Figure A4: Additional qualitative comparisons with camera-control baselines. SemCam (highlighted in blue) against Vista4D, TrajectoryCrafter, ReCamMaster, and CamCloneMaster on further scenes, one block per camera behavior; rows show the input and method outputs, columns four frames sampled over time.
Figure A5: Additional qualitative comparisons with camera-control baselines (continued). SemCam (highlighted in blue) against Vista4D, TrajectoryCrafter, ReCamMaster, and CamCloneMaster on further scenes, one block per camera behavior; rows show the input and method outputs, columns four frames sampled over time.
Figure A6: Screenshot of the Google Forms interface shown to participants in the user study.
Figure A7: Example of a user study question showing intended camera motion, reference video (Ref), generated video options (A and B), and the corresponding evaluation criteria (Camera-Control Alignment and Visual Alignment).
Motion
Subject appearance
VLM judge
Variant
Sem. motion ↑
Success ↑
DINO ↑
ID-Sim ↓
Sem. motion ↑
Identity ↑
Design components
Full (Ours)
0.654
0.706
0.636
0.283
0.610
0.737
− background loss
0.611
0.634
0.638
0.284
0.602
0.720
− motion modulation
0.530
0.554
0.662
0.259
0.551
0.768
Shared LoRA ( AB )
0.306
0.277
0.704
0.225
0.436
0.809
Appendix
Table A4: Full ablation on the 33-scene subset. Best per column over all variants in bold .