Unveiling the Value of Motion for Cinematic Camera Trajectories
Organizations: University of Edinburgh · The Hong Kong University of Science and Technology
Abstract
Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation. To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGEN, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed not only helps numerically to achieve better alignment and generation, but also inherently encodes complex directorial intent.
Figures & tables
| Model | Rep. | R@1 | R@5 | R@10 | MedR | AlignScore | #Params |
|---|---|---|---|---|---|---|---|
| Ours | DirSpeed | 25.2 | 40.7 | 49.0 | 11 | 66.3 | 3.6 M |
| Pose9D | 17.8 | 26.5 | 31.4 | 51 | 46.6 | 3.6 M | |
| CLaTr [ 9 ] | DirSpeed | 19.7 | 30.5 | 39.2 | 23 | 70.6 | 30 M |
| Pose9D | 6.9 | 12.0 | 14.7 | 361 | 37.4 | 30 M |
| Attribute | Task | #Class | DirSpeed | Pose9D | |
|---|---|---|---|---|---|
| Era | single-label | 3 | 49.8 | 47.3 | |
| Genre | multi-label | 3 | 62.1 | 60.7 | |
| Director | single-label | 3 | 68.2 | 51.8 |
| Trajectory Quality | Text–Trajectory Alignment | Movie Attributes | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method / Setting | F1 | FCD | Coverage | AlignScore | R@1 | MedR | Era | Genre | Director |
| CCD ∗ [ 8 ] | 0.117 | 52.45 | 0.298 | 4.90 | 0.08 | 1032.5 | 44.3 | 46.3 | 12.1 |
| CCD [ 8 ] | 0.128 | 49.19 | 0.270 | 18.82 | 0.43 | 397.0 | 47.2 | 56.3 | 32.7 |
| E.T. ∗ [ 9 ] | 0.015 | 135.74 | 0.034 | 0.00 | 0.08 | 960.0 | 40.7 | 56.7 | 25.0 |
| E.T. [ 9 ] | 0.000 | 179.88 | 0.014 | 0.36 | 0.16 | 745.0 | 40.7 | 45.8 | 33.6 |
| GenDoP ∗ [ 10 ] | 0.175 | 102.97 | 0.085 | 7.63 | 0.23 | 715.5 | 33.5 | 45.6 | 18.1 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Scale | Annotations | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | #Samples | #Frames | Avg(s) | motion | logline | movie attributes | Source |
| CCD [ 8 ] | 25 K | 4.5 M | 7.2 | ✓ | ✗ | ✗ | Synthetic |
| E.T. [ 9 ] | 115 K | 11 M | 3.8 | ✓ | ✗ | ✗ | Movie |
| DataDoP [ 10 ] | 29 K | 11 M | 14.4 | ✓ | ✗ | ✗ | Movie |
| CineScript (ours) | 28 K | 10 M | 12.0 | ✓ | ✓ | ✓ ( K) | Movie |
| Attribute | Classes | Type |
|---|---|---|
| Era | Film era ( ); Early digital ( – ); Digital mature ( ) | single-label |
| Genre | Drama/Romance; Comedy; Action/Thriller/Sci-Fi | multi-label |
| Director | Christopher Nolan; Wes Anderson; Steven Spielberg | single-label |
| Trajectory Quality | Text–Trajectory Alignment | Movie Attributes | |||||||
| Method | F1 | FCD | Coverage | AlignScore | R@1 | MedR | Era | Genre | Director |
| CCD ∗ [ 8 ] | 0.117 | 52.45 | 0.298 | 4.90 | 0.08 | 1032.5 | 44.3 | 46.3 | 12.1 |
| CCD [ 8 ] | 0.128 | 49.19 | 0.270 | 18.82 | 0.43 | 397.0 | 47.2 | 56.3 | 32.7 |
| CCD † [ 8 ] | 0.168 | 25.96 | 0.576 | 11.71 | 0.35 | 690.0 | 42.5 | 48.5 | 39.8 |
| E.T. ∗ [ 9 ] | 0.015 | 135.74 | 0.034 | 0.00 | 0.08 | 960.0 | 40.7 | 56.7 | 25.0 |
| E.T. [ 9 ] | 0.194 | 16.39 | 0.596 | 29.78 | 0.81 | 251.0 | 42.0 | 57.5 | 28.1 |
| Method | F1 | FCD | Coverage | AlignScore | R@1 | MedR | Era | Genre | Director |
|---|---|---|---|---|---|---|---|---|---|
| CineGEN | 0.437 | 6.77 | 0.783 | 57.79 | 3.35 | 52 | 47.8 | 60.7 | 48.1 |
| CineGEN caption-only | 0.419 | 6.76 | 0.749 | 65.83 | 5.82 | 42 | 45.6 | 60.6 | 41.6 |
| GenDoP [ 10 ] | 0.234 | 22.10 | 0.586 | 33.11 | 0.93 | 195 | 33.2 | 52.5 | 29.9 |
| Method | FCD | Coverage | R@1 | MedR |
|---|---|---|---|---|
| CCD ∗ [ 8 ] | 195.23 | 0.293 | 0.39 | 1158.5 |
| CCD [ 8 ] | 567.67 | 0.253 | 0.43 | 965.5 |
| E.T. ∗ [ 9 ] | 1311.94 | 0.006 | 0.04 | 1175.0 |
| E.T. [ 9 ] | 111.71 | 0.593 | 0.39 | 558.0 |
| GenDoP ∗ [ 10 ] | 339.73 | 0.188 | 0.12 | 810.0 |
| GenDoP [ 10 ] | 59.85 | 0.656 | 0.27 | 467.5 |
| Evaluator | Matched cosine | Mismatched cosine | R@1 | MedR |
|---|---|---|---|---|
| CLaTr- DirSpeed | 70.6 | 19.7 | 23 | |
| Ours- DirSpeed | 66.3 | 25.2 | 11 |
| Trajectory Quality | Text–Trajectory Alignment | Movie Attributes | |||||||
| Setting | F1 | FCD | Coverage | AlignScore | R@1 | MedR | Era | Genre | Director |
| CineGEN (ours) | 0.437 | 6.77 | 0.783 | 57.79 | 3.35 | 52.0 | 47.8 | 60.7 | 48.1 |
| Pose9D rep. | 0.189 | 20.99 | 0.638 | 41.11 | 1.94 | 110.5 | 42.5 | 58.2 | 29.8 |
| w/o | 0.425 | 15.13 | 0.722 | 53.07 | 3.84 | 54.0 | 40.2 | 61.2 | 35.2 |
| w/o | 0.402 | 9.75 | 0.769 | 51.16 | 2.44 | 72.0 | 47.7 | 56.1 | 45.1 |
| w/o variance | 0.401 | 8.84 | 0.784 | 53.60 | 2.79 | 58.0 | 42.1 | 61.2 | 43.2 |
| Feature | R@1 | R@5 | R@10 | MedR | AlignScore |
|---|---|---|---|---|---|
| direction-only | 19.7 | 33.0 | 40.1 | 22 | 61.2 |
| speed-only | 5.7 | 9.1 | 12.3 | 290 | 46.8 |
| velocity | 3.8 | 8.0 | 9.3 | 377 | 48.4 |
| DirSpeed (ours) | 25.2 | 40.7 | 49.0 | 11 | 66.3 |
| Rep. | Text | AlignScore | R@1 | R@5 | R@10 | MedR |
|---|---|---|---|---|---|---|
| DirSpeed | motion | 66.3 | 25.2 | 40.7 | 49.0 | 11 |
| motion + logline | 50.6 | 4.6 | 14.6 | 23.7 | 41 | |
| logline | 31.1 | 1.0 | 2.9 | 4.8 | 227 | |
| Pose9D | motion | 38.7 | 8.7 | 15.9 | 21.5 | 66 |
| motion + logline | 33.3 | 2.3 | 7.7 | 12.3 | 102 | |
| logline | 26.2 | 0.7 | 1.9 | 3.0 | 428 |
| Representation | AlignScore | R@1 | MedR |
|---|---|---|---|
| First-frame-canonical Pose9D | 46.4 | 15.2 | 51 |
| + Scale normalization Pose9D | 49.0 | 18.9 | 40 |
| DirSpeed | 66.3 | 25.2 | 11 |
| Magnitude parameterization | AlignScore | R@1 | MedR |
|---|---|---|---|
| Linear speed | 63.6 | 21.5 | 14 |
| Global z-score | 64.7 | 24.4 | 11 |
| 64.5 | 22.4 | 14 | |
| 66.3 | 25.2 | 11 |
| Representation | Short F1 | Long F1 | Short FCD | Long FCD | Short Cov. | Long Cov. |
|---|---|---|---|---|---|---|
| Pose9D | 0.240 | 0.187 | 10.14 | 33.08 | 0.794 | 0.635 |
| DirSpeed | 0.411 | 0.419 | 10.72 | 17.72 | 0.820 | 0.786 |
| Representation | Input dim. | AlignScore | R@1 | MedR |
|---|---|---|---|---|
| Pose9D | 9 | 46.6 | 17.8 | 51 |
| DirSpeed | 8 | 66.3 | 25.2 | 11 |
| Pose9D + DirSpeed | 17 | 66.9 | 27.0 | 9 |
| Training data | Pose9D R@1 | DirSpeed R@1 |
|---|---|---|
| 10% | 11.3 | 15.1 |
| 25% | 11.0 | 15.9 |
| 50% | 12.2 | 20.2 |
| 100% | 17.8 | 25.2 |
| Model | Trainable params | Pose9D R@1 | DirSpeed R@1 |
|---|---|---|---|
| Small | 0.53M | 16.3 | 24.2 |
| Base | 3.49M | 17.8 | 25.2 |
| Large | 26.14M | 17.1 | 25.0 |
| Hyperparameter | Value |
|---|---|
| Trajectory Encoder (Trainable, M params) | |
| Input dimension | 8 ( DirSpeed ) or 9 ( Pose9D ) |
| Linear projection | |
| Positional encoding | Sinusoidal (max length ) |
| Transformer layers | |
| Transformer dimensions | , heads, |
| Hyperparameter | Value |
|---|---|
| Conditioning and Architecture | |
| Text encoder | CLIP ViT-B/32 (frozen, dims) |
| Logline embedding dim | |
| Sequencer | TransformerAdaLN ( block) |
| Sequencer dimensions | , heads, , Dropout |
| Conditioning fusion | Text ( ) + First Pose ( ) + Logline ( ) = dims dims |