cs.CVOct 8, 2026

TKCAM: Text and Keyframe to Camera Trajectory Generation

Authors: Haozhe Yang, Zhiyang Dou, Zekai Gu, Cheng Lin, Wenping Wang, Yuan Liu, Taku Komura

Organizations: The University of Hong Kong · The Hong Kong University of Science and Technology · Macau University of Science and Technology · Texas A&M University

Abstract

Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Fréchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization. Code is available at https://github.com/linearalgebrayhz/TKCAM.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens

    Apr 21, 2026Xinxuan Lu, Charless Fowlkes, Alexander C. BergT2I GenerationControllable Image Generation

  2. ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation

    May 7, 2026Omar El Khalifi, Thomas Rossi, Oscar Fossey +6Diffusion Model GuidanceCamera-Controlled Video Generation

  3. TriMotion: Modality-Agnostic Camera Control for Video Generation

    Jun 18, 2026Seunghyun Shin, Jifei Song, Wooseok Jeon +2Camera-Controlled Video GenerationMultimodal Learning