AESOP: Asymmetric Human-Camera Generation with Translation-Intensity Control
Organizations: East China Normal University · Peking University · Sun Yat-sen University · Tencent
Abstract
Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.
Figures & tables
| camera distribution & alignment | camera framing & geometry | human distribution & alignment | ||||||||
| Method | FDC | CLaTr | F1 | Recall C | r-FPD | Out | ADE | Rot. | FD TMR | TMR |
| GT reference | 70.237 | 0.945 | 1.00 | 0.003 | 0.71 | 0.000 | 0.002 | 18.398 | ||
| PulpMotion AE (sym.) | 16.59 3.52 | 59.91 1.38 | 0.739 0.032 | 0.973 0.005 | 0.133 0.02 | 3.93 0.55 | 0.110 0.01 | 1.37 0.10 | 38.79 9.99 | 20.15 0.28 |
| AESOP AE (asym.) | 0.11 0.10 | 70.10 0.13 | 0.941 0.001 | 0.999 0.0004 | 0.089 0.01 | 2.93 0.28 | 0.029 0.0006 | 0.41 0.01 | 10.05 0.89 | 17.56 0.12 |
| Given-human camera generation | ||||||||||
| DanceCamera3D | 180.08 5.24 | 20.92 0.54 | 0.252 0.002 | 0.600 0.01 | 2.692 0.48 | 21.72 2.33 | 2.995 0.07 | 63.52 1.19 | – | – |
| Method | Precision | Recall | Density | Coverage |
|---|---|---|---|---|
| PulpMotion DiT | 0.809 0.02 | 0.215 0.02 | 0.748 0.05 | 0.465 0.02 |
| PulpMotion MAR | 0.802 0.005 | 0.312 0.01 | 0.705 0.02 | 0.504 0.002 |
| AESOP | 0.879 0.002 | 0.835 0.01 | 0.933 0.02 | 0.785 0.01 |
| Compared with | Pref. | 95% CI |
|---|---|---|
| Given-human | ||
| CCD | 75.3 | [64.8, 85.5] |
| DanceCamera3D | 99.2 | [97.5, 100.0] |
| DIRECTOR-C | 86.6 | [81.5, 91.5] |
| Joint | ||
| PulpMotion DiT | 86.3 | [80.1, 91.9] |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Training | Reconstruction | ||
|---|---|---|---|
| Method | Independent human learning | camera-independent human | human-aware camera |
| PulpMotion (symmetric) | Both streams trained together | ||
| AESOP (asymmetric) | Train human first, then freeze | ||
| human reconstruction | camera reconstruction | |||||||
|---|---|---|---|---|---|---|---|---|
| Autoencoder | Global MPJPE | Aligned MPJPE | Root ADE | Root FDE | FDE | Trans. vel. | Rot. vel. | Trans. jerk |
| PulpMotion (sym.) | 0.1620 0.02 | 0.0617 0.01 | 0.1376 0.01 | 0.3283 0.04 | 0.1677 0.02 | 124.28 3.02 | 10.55 0.10 | 154.67 1.54 |
| AESOP (asym.) | 0.1304 0.01 | 0.0479 0.003 | 0.1063 0.01 | 0.2611 0.02 | 0.0356 0.002 | 10.11 0.81 | 10.98 0.21 | 8.55 0.40 |
| Stage | Phase | Updates | Batch | Learning rate | Frozen components |
|---|---|---|---|---|---|
| Reconstruction | human autoencoder | 210K | 128 | None | |
| camera autoencoder | 210K | 128 | human autoencoder | ||
| Generation | human flow | 105K | 128 | Autoencoders | |
| Original camera | 105K | 128 | Autoencoders, human flow | ||
| camera + DPA | 35K | 120 | Autoencoders, human flow | ||
| camera + DPA + IPA | 15K | 120 | Autoencoders, human flow |
| Operation | Acceptance rules |
|---|---|
| DPA: direction pairs | |
| Select source events | Duration frames; same-sign velocity fraction ; signed net displacement/path . Truck displacement m; Dolly displacement . |
| Check target geometry | Opposite-target human distance m; opposite/original response . Rotation/FOV unchanged; Truck vertical shift m and radial relative change ; Dolly radial-direction change . |
| Check framing and effect | No zero-visible frame; visibility ; opposite outscreen increase . Truck: projected center shift image width or viewpoint change . Dolly: peak box-height ratio change . |
| Check reconstructed targets | Feature MSE ; both direction signs correct; decoded/raw response ratios . No framing failure; and . |
| IPA: intensity pairs | |
| Method | Inputs and conditioning | Sampling and guidance |
|---|---|---|
| DIRECTOR-C | Camera text and human root trajectory | EDM: 10 Euler/Heun steps; camera CFG 1.4 |
| DanceCamera3D | Camera text and full human motion; text replaces music | DDIM: 50 steps, ; human/camera CFG 1.75/1 |
| CCD | Camera text; text-only network | DDPM: 1,000 steps; camera CFG 2 |
| CCD-H | Camera text and human joints; text cross-attention | DDPM: 1,000 steps; camera CFG 2 |
| PulpMotion DiT | Human and camera text; symmetric representation | DDPM: 50 steps; |
| PulpMotion MAR | Human and camera text; symmetric representation | 18 autoregressive iterations, each with 50 DDPM steps; |
| Given-human generation | Joint generation | |||||
|---|---|---|---|---|---|---|
| Checkpoint | Original | Reversed | Combined | Original | Reversed | Combined |
| Camera original | 86.38 0.86 | 75.66 0.68 | 81.02 0.15 | 83.15 0.13 | 76.42 0.43 | 79.78 0.28 |
| 35K factual-pair control | 87.60 0.43 | 77.71 0.13 | 82.66 0.17 | 84.95 1.13 | 78.51 0.17 | 81.73 0.51 |
| + DPA | 89.66 0.67 | 79.00 0.22 | 84.33 0.28 | 86.63 0.77 | 79.58 0.50 | 83.11 0.14 |