Organizations: The University of Hong Kong · The Hong Kong University of Science and Technology · Macau University of Science and Technology · Texas A&M University
Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Fréchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization. Code is available at https://github.com/linearalgebrayhz/TKCAM.
Figures & tables
Figure 1: Motion distribution comparison across camera trajectory datasets. Left: Coarse-grained breakdown into static, single-axis, and compound motion (DataDoP and E.T. values estimated from their published figures). Right: Top-13 motion combinations in RealEstate10K-Cap; the remaining 35.1% is distributed across 50+ other combination patterns.
Dataset
#Samples
Domain
Caption
Vocab
Avg. Length (s)
RealEstate10K [ 38 ]
79K
Indoor
×
-
-
CCD [ 16 ]
25K
Synthetic
Synthetic
48
7.2
E.T. [ 7 ]
115K
Film
Camera-Char
1.7K
3.8
DataDoP [ 36 ]
29K
Film
Directorial
8.7K
14.4
Ours
25K
Indoor
Scene-aware
3.8K
6.72
Table 1: Comparison with existing camera trajectory datasets. Our dataset focuses on indoor environments with scene-aware captions.
Figure 2: Overview of our proposed pipeline. Given per-frame camera poses M as trajectory, a text description T and sparse keyframe images V , we first encode the continuous 12-dimensional camera trajectory into latent features and discretize them with a multi-level Residual Vector Quantizer (RVQ), producing a hierarchy of base-layer (green) and residual-layer (orange) motion tokens. During training, a masked base transformer predicts masked base-layer tokens conditioned on the text embedding from a frozen T5 encoder and visual tokens extracted by a frozen CLIP encoder, while a masked residual transformer refines higher-level residual tokens conditioned on the previously predicted tokens and the same multimodal context. Finally, a motion decoder reconstructs smooth and temporally coherent camera trajectories from the predicted hierarchical motion tokens.
Method
FID ↓
Matching ↑
R@1 ↑
R@3 ↑
R@10 ↑
Δ Div ↓
Track A: Text-Conditioned (1500 samples)
CCD [ 16 ]
0.732
0.022
0.07
0.27
1.20
0.156
CCD (fine-tuned) [ 16 ]
0.594
0.037
0.33
0.67
2.33
0.075
E.T. [ 7 ]
1.135
0.000
0.07
0.27
0.67
0.538
E.T. (fine-tuned) [ 7 ]
0.559
0.050
0.40
1.13
2.93
0.098
Director3D [ 19 ]
0.831
0.005
0.13
0.33
0.93
0.164
Table 2: Quantitative evaluation of text-to-camera trajectory generation on the mixed benchmark. Metrics include Fréchet distance computed on Universal CLaTr features (FID), text-trajectory Matching Score, and Retrieval metrics (R@ 1 , R@ 3 , R@ 10 , in %) computed over the full evaluation set. Δ Div represents the absolute difference between the generated trajectory diversity and the reference set diversity (closer to 0 is better). “Fine-tuned” denotes models fine-tuned/retrained on RealEstate10K-Cap. Bold indicates the best performance within each track.
Figure 3: Qualitative comparison of generations. While baselines struggle with motion magnitude, directional alignment, or trajectory quality (e.g., CCD’s circular loops), TKCAM produces semantically faithful paths that capture complex compound motions.
Figure 4: TKCAM as a geometric motion prior for downstream video generation tasks.
RVQ Stages
FID ↓
Match. ↑
R@1 ↑
R@3 ↑
R@10 ↑
Δ Div ↓
L1 (Base Only)
0.582
0.094
1.40
4.07
9.33
0.081
L1 + L2
0.522
0.102
0.93
4.07
9.20
0.068
L1 + L2 + L3
0.525
0.099
1.33
4.27
8.73
0.063
Full (L1–L4)
0.529
0.100
1.13
3.80
9.60
0.065
Table 3: Ablation on hierarchical RVQ levels.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Complete motion category distribution of RealEstate10K-Cap, extending Figure 1 to all identified patterns. Categories are sorted by frequency and color-coded by motion family.
Module
Hyperparameter
Value
RVQ Autoencoder
Codebook Size ( V )
256
Quantization Levels ( K )
4
Quantizer Dropout Prob ( p )
0.2
Kinematic Loss Weight ( λexp )
0.5
Orthogonality Weight ( λorth )
0.1
Commitment Weight ( β )
0.02
Appendix
Table 4: Detailed hyperparameters for TKCAM training.
Test Domain
FID ↓
Δ Div ↓
Re10K-Cap Train
DataDoP Train
Re10K-Cap Train
DataDoP Train
Mixed
0.967
1.190
0.196
0.389
RealEstate10K
1.204
1.458
0.274
0.394
E.T. (OOD)
0.956
1.134
0.013
0.179
DataDoP
1.431
1.572
0.134
0.466
Appendix
Table 5: Controlled comparison between RealEstate10K-Cap and DataDoP as training sources. The same TKCAM architecture is trained on equally sized subsets of the two datasets. Lower FID and Δ Div are better.
Method
Trans. Acc.
Trans. Jerk
Ang. Vel.
Ang. Acc.
Ang. Jerk
Mean per-frame derivative magnitude
GT (reference)
0.00078
0.00054
0.00460
0.00134
0.00186
TKCAM (Sparse Keyframes)
0.00045
0.00024
0.00745
0.00718
0.01187
TKCAM (First-Frame)
0.00032
0.00016
0.00629
0.00641
0.01067
GenDoP (RGBD)
0.00092
0.00133
0.00394
0.00042
0.00058
1D Wasserstein distance to GT ↓
Appendix
Table 6: Temporal smoothness analysis in the vision-conditioned setting. Top: mean per-frame derivative magnitudes. Bottom: 1D Wasserstein distance between each derivative distribution and ground truth (lower is closer to real camera dynamics).
Figure 6: Additional qualitative comparisons with baselines.
Figure 7: Two representative failure cases of TKCAM: tilt underestimation and action omission in multi-stage sequences.