cs.CVSep 29, 2026

AESOP: Asymmetric Human-Camera Generation with Translation-Intensity Control

Authors: Jingzhong Lin, Zhanke Wang, Heng Li, Wenxiang Liu, Zhao Zhang, Kecheng Tang, Dongdong Xiang, Changbo Wang, +4 more

Organizations: East China Normal University · Peking University · Sun Yat-sen University · Tencent

Abstract

Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

    Jun 1, 2026Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra +4Cinematic IdealDramadirector

  2. UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

    Aug 3, 2026Liming Tan, Ye Chen, Hao Zhang +3Camera ControlHuman Motion Generation

  3. ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation

    May 7, 2026Omar El Khalifi, Thomas Rossi, Oscar Fossey +6Camera-Conditioned Video GenerationVideo Generation