Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.
Figures & tables
Figure 1: AESOP supports (a) camera generation for a given human motion, (b) joint human–camera generation, and (c) continuous camera translation-intensity control (the amount of camera travel) in both tasks without changing the supplied or generated human motion.
Figure 2: AESOP overview. (A) Separate encoders produce 128-channel human and 64-channel camera latents; the camera decoder reads their concatenation. (B,C) Human generation supplies the complete latent sequence to both human decoding and camera conditioning. Given-human generation instead encodes the supplied motion. (D,E) Human and camera Transformer blocks.
Figure 3: Training target construction of DPA and IPA. (a,b) Opposite directions share human motion and the initial camera state. (c,d) Intensity variants retain human motion and camera text.
camera distribution & alignment
camera framing & geometry
human distribution & alignment
Method
FDC ↓
CLaTr ↑
F1 ↑
Recall C ↑
r-FPD ↓
Out ↓
ADE ↓
Rot. ↓
FD TMR ↓
TMR ↑
GT reference
≈0
70.237
0.945
1.00
0.003
0.71
0.000
0.002
≈0
18.398
PulpMotion AE (sym.)
16.59 ± 3.52
59.91 ± 1.38
0.739 ± 0.032
0.973 ± 0.005
0.133 ± 0.02
3.93 ± 0.55
0.110 ± 0.01
1.37 ± 0.10
38.79 ± 9.99
20.15 ± 0.28
AESOP AE (asym.)
0.11 ± 0.10
70.10 ± 0.13
0.941 ± 0.001
0.999 ± 0.0004
0.089 ± 0.01
2.93 ± 0.28
0.029 ± 0.0006
0.41 ± 0.01
10.05 ± 0.89
17.56 ± 0.12
Given-human camera generation
DanceCamera3D
180.08 ± 5.24
20.92 ± 0.54
0.252 ± 0.002
0.600 ± 0.01
2.692 ± 0.48
21.72 ± 2.33
2.995 ± 0.07
63.52 ± 1.19
–
–
Table 1: Quantitative comparison. AE rows report reconstruction; the remaining model rows report generation. AESOP (no aug.) uses original pairs, (+DPA) adds direction pairs, and full AESOP further adds intensity pairs. Human scores are shared across AESOP camera-training phases. Bold marks the best displayed value within each reconstruction/task group, excluding GT; blue marks AESOP.
Figure 4: Qualitative comparison on test examples. (a) Given-human camera generation with a pull-out prompt. (b) Joint human–camera generation with a push-in prompt. Each method has a global view and three selected camera projections, ordered from top to bottom in time. Global views show human–camera geometry at individually fitted scales; projections show framing over time.
Method
Precision ↑
Recall ↑
Density ↑
Coverage ↑
PulpMotion DiT
0.809 ± 0.02
0.215 ± 0.02
0.748 ± 0.05
0.465 ± 0.02
PulpMotion MAR
0.802 ± 0.005
0.312 ± 0.01
0.705 ± 0.02
0.504 ± 0.002
AESOP
0.879 ± 0.002
0.835 ± 0.01
0.933 ± 0.02
0.785 ± 0.01
Table 2: Human distributional fidelity and coverage in joint generation in the common frozen human embedding space. Bold marks the best value in each column.
Figure 5: Translation-intensity examples with shared human context, text and camera noise. Outputs are regenerated at each intensity; numbers report first-to-last camera-center displacement. Fading snapshots indicate temporal order. Global views are fitted separately to each trajectory's spatial extent.
Compared with
Pref.
95% CI
Given-human
CCD
75.3
[64.8, 85.5]
DanceCamera3D
99.2
[97.5, 100.0]
DIRECTOR-C
86.6
[81.5, 91.5]
Joint
PulpMotion DiT
86.3
[80.1, 91.9]
Table 3: Overall-quality preference for AESOP (%). Scores average clip-level responses, with ties receiving half credit. Intervals are participant-cluster 95% CIs.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: human–camera representation dependencies. (a) PulpMotion’s symmetric representation encodes both inputs with a shared encoder, reconstructs each from its latent, and learns an auxiliary framing latent from their concatenation ( Courant et al., 2026 ) . (b) AESOP separately encodes the two inputs. The human decoder reads only human latents; the camera decoder concatenates camera latents with frozen human latents (circled C). Human representation training is completed before camera training.
Training
Reconstruction
Method
Independent human learning
camera-independent human
human-aware camera
PulpMotion (symmetric)
× Both streams trained together
×
✓
AESOP (asymmetric)
✓ Train human first, then freeze
✓
✓
Appendix
Table S1: Training and reconstruction dependencies. Both representations provide human context for camera reconstruction. AESOP additionally separates human learning from camera training and makes human reconstruction independent of camera input.
human reconstruction
camera reconstruction
Autoencoder
Global MPJPE ↓
Aligned MPJPE ↓
Root ADE ↓
Root FDE ↓
FDE ↓
Trans. vel. ↓
Rot. vel. ↓
Trans. jerk ↓
PulpMotion (sym.)
0.1620 ± 0.02
0.0617 ± 0.01
0.1376 ± 0.01
0.3283 ± 0.04
0.1677 ± 0.02
124.28 ± 3.02
10.55 ± 0.10
154.67 ± 1.54
AESOP (asym.)
0.1304 ± 0.01
0.0479 ± 0.003
0.1063 ± 0.01
0.2611 ± 0.02
0.0356 ± 0.002
10.11 ± 0.81
10.98 ± 0.21
8.55 ± 0.40
Appendix
Table S2: Autoencoder reconstruction on the test split. Sym./asym. denote PulpMotion's symmetric and AESOP's asymmetric autoencoders. Motion errors are clip-balanced and measured in mm/s, ∘ /s and m/s 3 , respectively. MPJPE, ADE and FDE are in meters; aligned MPJPE removes root translation. Bold marks the best value between the autoencoders.
Stage
Phase
Updates
Batch
Learning rate
Frozen components
Reconstruction
human autoencoder
210K
128
5×10−5
None
camera autoencoder
210K
128
5×10−5
human autoencoder
Generation
human flow
105K
128
2×10−4
Autoencoders
Original camera
105K
128
10−4
Autoencoders, human flow
camera + DPA
35K
120
2×10−5
Autoencoders, human flow
camera + DPA + IPA
15K
120
2×10−5
Autoencoders, human flow
Appendix
Table S3: Training schedule. The final phase additionally trains the intensity MLP at 10−4 .
Figure S2: DPA construction and validation. Left: the original Dolly-in trajectory. Middle: its opposite Dolly-out target, with human motion, camera rotation, field of view and timing fixed. Right: a reconstruction rejected because its direction change exceeds the acceptance threshold after decoding. The gold point marks the initial camera center; global views show the starting human pose and full trajectory, and projections show the synchronized final frame. The three stages summarize source filtering, direction reversal and target checks; 601 sources pass the final reconstruction check.
Figure S3: IPA target construction. Reduced, original and increased translation share human motion and camera text. Camera travel changes from 0.356 to 0.634 to 0.786 m in this example. The gold point marks the initial camera center; global views and synchronized final-frame projections show the resulting geometry and framing. The three stages summarize the active branch, which supplies 16,000 training targets from 8,000 sources. The accompanying null branch supplies 2,000 low-translation sources with unchanged targets.
Operation
Acceptance rules
DPA: direction pairs
Select source events
Duration ≥45 frames; same-sign velocity fraction ≥0.80 ; signed net displacement/path ≥0.60 . Truck displacement ≥0.20 m; Dolly displacement ≥max(0.20 m,0.10ρts) .
Check target geometry
Opposite-target human distance >0.25 m; opposite/original response ∈[0.80,1.25] . Rotation/FOV unchanged; Truck vertical shift ≤10−5 m and radial relative change ≤0.10 ; Dolly radial-direction change ≤0.05∘ .
Check framing and effect
No zero-visible frame; visibility ≥0.80 ; opposite outscreen increase ≤0.05 . Truck: projected center shift ≥0.08 image width or viewpoint change ≥6∘ . Dolly: peak box-height ratio change ≥0.12 .
Check reconstructed targets
Feature MSE ≤0.01181 ; both direction signs correct; decoded/raw response ratios ∈[0.70,1.30] . No framing failure; Vdec≥1 and Vdec/Vraw≥0.70 .
IPA: intensity pairs
Appendix
Table S4: Construction acceptance criteria for DPA and IPA, grouped by source selection, target geometry, framing and reconstruction.
Method
Inputs and conditioning
Sampling and guidance
DIRECTOR-C
Camera text and human root trajectory
EDM: 10 Euler/Heun steps; camera CFG 1.4
DanceCamera3D
Camera text and full human motion; text replaces music
DDIM: 50 steps, η=1 ; human/camera CFG 1.75/1
CCD
Camera text; text-only network
DDPM: 1,000 steps; camera CFG 2
CCD-H
Camera text and human joints; text cross-attention
DDPM: 1,000 steps; camera CFG 2
PulpMotion DiT
Human and camera text; symmetric representation
DDPM: 50 steps; (gc,gm,gz)=(11,−1,0)
PulpMotion MAR
Human and camera text; symmetric representation
18 autoregressive iterations, each with 50 DDPM steps; (gc,gm,gz)=(3.5,2,0)
Appendix
Table S5: Baseline inputs, conditioning and sampling. For PulpMotion, the tuple denotes conditional-text guidance, autoregressive-context guidance, and auxiliary projection guidance, respectively. A negative autoregressive-context weight selects standard CFG; zero auxiliary projection weight disables that guidance term. MAR generates autoregressively, with a diffusion sampler inside each iteration. EDM, DDIM and DDPM follow Karras et al. (2022) , Song et al. (2021) and Ho et al. (2020) , respectively.
Given-human generation
Joint generation
Checkpoint
Original ↑
Reversed ↑
Combined ↑
Original ↑
Reversed ↑
Combined ↑
Camera original
86.38 ± 0.86
75.66 ± 0.68
81.02 ± 0.15
83.15 ± 0.13
76.42 ± 0.43
79.78 ± 0.28
35K factual-pair control
87.60 ± 0.43
77.71 ± 0.13
82.66 ± 0.17
84.95 ± 1.13
78.51 ± 0.17
81.73 ± 0.51
+ DPA
89.66 ± 0.67
79.00 ± 0.22
84.33 ± 0.28
86.63 ± 0.77
79.58 ± 0.50
83.11 ± 0.14
Appendix
Table S6: Direction accuracy (%). DPA and factual continuations share the starting checkpoint, 35K updates, batch allocation and sampling schedule; factual pairs replace directional counterfactuals in the control. Bold marks the best mean in each column. Section D.2 defines paired sampling and scoring.
Figure S4: User-study interface. Anonymized A/B videos show the generated camera view above an external spatial view. Participants compare text alignment, motion quality, framing and task-specific overall quality. The joint task additionally evaluates human motion.
Figure S6: Camera generation conditioned on HumanML3D motions. Each example shows a global view and three successive camera views; lighter poses indicate earlier frames. Global views are fitted separately to each scene.
Figure S7: Joint generation from HumanML3D action descriptions. Each example shows generated human motion and camera trajectory in a global view and three successive camera views.
Figure S8: Camera generation conditioned on HY-Motion outputs. Four synthesized human motions are paired with static, pull-out and pull-in camera prompts. Each example includes a global view and three successive camera views.
Figure S9: Additional given-human camera generation examples. Global views are fitted separately to each trajectory.
Figure S10: Additional joint human–camera generation examples. Global views are fitted separately to each trajectory.
Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spatially inconsistent, and indifferent to the human subject driving the scene. In this work, we present Auteur, a method for language-driven, human-centric camera framing in generative video. Our core insight is that professional filmmakers conceive shots not as world-space trajectories but as framings defined relative to the actor, encoding shot size, angle, and composition as functions of human pose and motion. We formalize this intuition as a human-centric camera parameterization and introduce a Domain-Specific Language (DSL) that is convertible to standard 6-DoF camera parameters. A fine-tuned multimodal large language model then acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are deterministically interpolated into continuous camera trajectories, which are then provided as input to video generators. We train and evaluate Auteur on a new dataset of 34K aligned text, human motion, and DSL-annotated camera trajectories drawn from procedural synthesis and real-world movie footage from the CondensedMovies dataset. Auteur enables cinematographic framing of human-centered scenes, a capability largely absent in prior generative models. To assess this behavior, we propose new framing-focused metrics, and our experiments show that Auteur consistently outperforms existing methods. Project page is https://cyberiada.github.io/Auteur/
Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra +4
Koç University · University College London, Adobe · Adobe +1
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typically rely on visual motion sequences, such as skeleton maps, pose maps, or rendered body representations, for motion control, while using camera embeddings for camera control. Such heterogeneous control interfaces force video generation models to reconcile pixel-aligned visual cues with non-visual geometric embeddings, making motion-camera attribution difficult and sensitive to camera estimation errors. We propose \textbf{UniMoCa}, a representation-driven framework that unifies motion and camera controls in visual space. At the core of UniMoCa is \textbf{Motion-Camera Visual Proxy} (\textbf{MCVP}), a mutually-sharable novel representation that converts 3D human motion and camera trajectories extracted from driving videos into an identity-neutral visual proxy. MCVP renders temporally aligned human geometry under the recovered camera trajectory and augments it with explicit camera trajectory markers, replacing heterogeneous visual-parametric controls with distinguishable visual cues. As both control factors are represented in the same visual space, they become mutually compatible rather than heterogeneous, enabling consistent joint reasoning and editing during video generation. We further curate a \textbf{MCVP-Video} dataset covering complex actions, multi-person interactions, and diverse camera trajectories. Experiments based on the Wan2.2 I2V show that UniMoCa achieves substantial gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity. More details are shown in our Project page: https://tanliming-daniel.github.io/UniMoCa/.
Liming Tan, Ye Chen, Hao Zhang +3
Shanghai Jiao Tong University · Monash University · USC-SJTU Institute of Cultural and Creative Industry
For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.
Omar El Khalifi, Thomas Rossi, Oscar Fossey +6
Kinetix, France · University of Oxford, United Kingdom · MBZUAI, United Arab Emirates