Organizations: The University of Hong Kong · The Hong Kong University of Science and Technology · Macau University of Science and Technology · Texas A&M University
Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Fréchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization. Code is available at https://github.com/linearalgebrayhz/TKCAM.
Figures & tables
Figure 1: Motion distribution comparison across camera trajectory datasets. Left: Coarse-grained breakdown into static, single-axis, and compound motion (DataDoP and E.T. values estimated from their published figures). Right: Top-13 motion combinations in RealEstate10K-Cap; the remaining 35.1% is distributed across 50+ other combination patterns.
Dataset
#Samples
Domain
Caption
Vocab
Avg. Length (s)
RealEstate10K [ 38 ]
79K
Indoor
×
-
-
CCD [ 16 ]
25K
Synthetic
Synthetic
48
7.2
E.T. [ 7 ]
115K
Film
Camera-Char
1.7K
3.8
DataDoP [ 36 ]
29K
Film
Directorial
8.7K
14.4
Ours
25K
Indoor
Scene-aware
3.8K
6.72
Table 1: Comparison with existing camera trajectory datasets. Our dataset focuses on indoor environments with scene-aware captions.
Figure 2: Overview of our proposed pipeline. Given per-frame camera poses M as trajectory, a text description T and sparse keyframe images V , we first encode the continuous 12-dimensional camera trajectory into latent features and discretize them with a multi-level Residual Vector Quantizer (RVQ), producing a hierarchy of base-layer (green) and residual-layer (orange) motion tokens. During training, a masked base transformer predicts masked base-layer tokens conditioned on the text embedding from a frozen T5 encoder and visual tokens extracted by a frozen CLIP encoder, while a masked residual transformer refines higher-level residual tokens conditioned on the previously predicted tokens and the same multimodal context. Finally, a motion decoder reconstructs smooth and temporally coherent camera trajectories from the predicted hierarchical motion tokens.
Method
FID ↓
Matching ↑
R@1 ↑
R@3 ↑
R@10 ↑
Δ Div ↓
Track A: Text-Conditioned (1500 samples)
CCD [ 16 ]
0.732
0.022
0.07
0.27
1.20
0.156
CCD (fine-tuned) [ 16 ]
0.594
0.037
0.33
0.67
2.33
0.075
E.T. [ 7 ]
1.135
0.000
0.07
0.27
0.67
0.538
E.T. (fine-tuned) [ 7 ]
0.559
0.050
0.40
1.13
2.93
0.098
Director3D [ 19 ]
0.831
0.005
0.13
0.33
0.93
0.164
Table 2: Quantitative evaluation of text-to-camera trajectory generation on the mixed benchmark. Metrics include Fréchet distance computed on Universal CLaTr features (FID), text-trajectory Matching Score, and Retrieval metrics (R@ 1 , R@ 3 , R@ 10 , in %) computed over the full evaluation set. Δ Div represents the absolute difference between the generated trajectory diversity and the reference set diversity (closer to 0 is better). “Fine-tuned” denotes models fine-tuned/retrained on RealEstate10K-Cap. Bold indicates the best performance within each track.
Figure 3: Qualitative comparison of generations. While baselines struggle with motion magnitude, directional alignment, or trajectory quality (e.g., CCD’s circular loops), TKCAM produces semantically faithful paths that capture complex compound motions.
Figure 4: TKCAM as a geometric motion prior for downstream video generation tasks.
RVQ Stages
FID ↓
Match. ↑
R@1 ↑
R@3 ↑
R@10 ↑
Δ Div ↓
L1 (Base Only)
0.582
0.094
1.40
4.07
9.33
0.081
L1 + L2
0.522
0.102
0.93
4.07
9.20
0.068
L1 + L2 + L3
0.525
0.099
1.33
4.27
8.73
0.063
Full (L1–L4)
0.529
0.100
1.13
3.80
9.60
0.065
Table 3: Ablation on hierarchical RVQ levels.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Complete motion category distribution of RealEstate10K-Cap, extending Figure 1 to all identified patterns. Categories are sorted by frequency and color-coded by motion family.
Module
Hyperparameter
Value
RVQ Autoencoder
Codebook Size ( V )
256
Quantization Levels ( K )
4
Quantizer Dropout Prob ( p )
0.2
Kinematic Loss Weight ( λexp )
0.5
Orthogonality Weight ( λorth )
0.1
Commitment Weight ( β )
0.02
Appendix
Table 4: Detailed hyperparameters for TKCAM training.
Test Domain
FID ↓
Δ Div ↓
Re10K-Cap Train
DataDoP Train
Re10K-Cap Train
DataDoP Train
Mixed
0.967
1.190
0.196
0.389
RealEstate10K
1.204
1.458
0.274
0.394
E.T. (OOD)
0.956
1.134
0.013
0.179
DataDoP
1.431
1.572
0.134
0.466
Appendix
Table 5: Controlled comparison between RealEstate10K-Cap and DataDoP as training sources. The same TKCAM architecture is trained on equally sized subsets of the two datasets. Lower FID and Δ Div are better.
Method
Trans. Acc.
Trans. Jerk
Ang. Vel.
Ang. Acc.
Ang. Jerk
Mean per-frame derivative magnitude
GT (reference)
0.00078
0.00054
0.00460
0.00134
0.00186
TKCAM (Sparse Keyframes)
0.00045
0.00024
0.00745
0.00718
0.01187
TKCAM (First-Frame)
0.00032
0.00016
0.00629
0.00641
0.01067
GenDoP (RGBD)
0.00092
0.00133
0.00394
0.00042
0.00058
1D Wasserstein distance to GT ↓
Appendix
Table 6: Temporal smoothness analysis in the vision-conditioned setting. Top: mean per-frame derivative magnitudes. Bottom: 1D Wasserstein distance between each derivative distribution and ground truth (lower is closer to real camera dynamics).
Figure 6: Additional qualitative comparisons with baselines.
Figure 7: Two representative failure cases of TKCAM: tilt underestimation and action omission in multi-stage sequences.
Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global scene understanding in text-to-image generation by learning parametric camera tokens. We fine-tune image generation models for viewpoint-conditioned text-to-image generation on a curated dataset that combines 3D-rendered images for geometric supervision and photorealistic augmentations for appearance and background diversity. Qualitative and quantitative experiments demonstrate that our method achieves state-of-the-art accuracy while preserving image quality and prompt fidelity. Unlike prior methods that overfit to object-specific appearance correlations, our viewpoint tokens learn factorized geometric representations that transfer to unseen object categories. Our work shows that text-vision latent spaces can be endowed with explicit 3D camera structure, offering a pathway toward geometrically-aware prompts for text-to-image generation. Project page: https://randdl.github.io/viewtoken_control/
For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.
Omar El Khalifi, Thomas Rossi, Oscar Fossey +6
Kinetix, France · University of Oxford, United Kingdom · MBZUAI, United Arab Emirates
Camera motion control is essential for directing viewpoint changes in generative systems. However, existing methods typically condition the generation process on a single specific modality, such as explicit pose trajectories or reference videos, limiting their ability to support heterogeneous user inputs. To address this limitation, we present TriMotion, a modality-agnostic framework for camera-controlled video generation that maps video, pose, and text inputs, describing the same camera trajectory into a shared motion embedding space. Learning such a space requires synchronized supervision across modalities. Therefore, we build the Motion Triplet Dataset by extending a Multi-Cam Video Dataset with geometry-grounded motion descriptions derived from camera extrinsics. We further introduce a latent motion consistency objective that leverages the motion embedding space to encourage the generated video to follow the target camera trajectory directly in latent space, avoiding the cost of pixel-space decoding. Extensive experiments show that TriMotion generates high-quality videos that accurately follow the target camera trajectories across all three modalities. Beyond standard generation, the shared motion embedding space also enables flexible applications such as sequential motion composition and cross-modal motion interpolation.
Seunghyun Shin, Jifei Song, Wooseok Jeon +2
GIST · Huawei Noah’s Ark Lab · Yonsei University +1