cs.CVOct 4, 2026

SemCam: Semantic Camera Motion Control for Video Generation

Authors: Janna Bruner, Omer Talmi, Ianir Ideses, Lior Fritz, Lior Wolf, Sagie Benaim

Organizations: Amazon Prime Video · Tel Aviv University · Hebrew University of Jerusalem

Abstract

Controlling the camera relative to a moving subject in an existing video is challenging: behaviors such as maintaining a frontal view require the camera to adapt to the subject's changing position and orientation, making the desired trajectory difficult to specify in advance. Existing camera-controlled video-to-video methods typically rely on explicit trajectories or reference motions, which do not directly express these dynamic camera--subject relationships. We introduce semantic camera motion control, a novel video-to-video task in which a reference video and a target motion label specify the desired subject-relative camera behavior without an explicit target trajectory. Our method, SemCam, learns to realize this behavior while preserving source content. It combines shared-basis low-rank adaptation with motion-conditioned modulation, while a background-consistency loss encourages fidelity in regions visible in both reference and target videos. We construct 661 paired videos covering eight semantic camera behaviors and evaluate on a separate 109-scene benchmark using subject-relative motion metrics, appearance measures, and a user study. SemCam achieves a semantic-motion success rate of 68.6%, compared with 45.3% for Vista4D, the strongest evaluated baseline, while maintaining comparable subject identity preservation.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TriMotion: Modality-Agnostic Camera Control for Video Generation

    Jun 18, 2026Seunghyun Shin, Jifei Song, Wooseok Jeon +2Camera ControlCamera Trajectories

  2. ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation

    May 7, 2026Omar El Khalifi, Thomas Rossi, Oscar Fossey +6Camera-Conditioned Video GenerationVideo Generation

  3. UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

    Aug 3, 2026Liming Tan, Ye Chen, Hao Zhang +3Camera ControlHuman Motion Generation