cs.CVSep 28, 2026

DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time

Authors: Ziqi Ma, Hongqiao Chen, Georgia Gkioxari

Abstract

Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation. Project website: https://glab-caltech.github.io/dynatokens/

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

    Sep 30, 2026Mikhail Dereviannykh, Vikram Voleti, Simon Donne +3Visual TokenizersAutoregressive Video Generation

  2. Wonder: Video World Model Done Better

    Jul 28, 2026Jiacong Xu, Hanwen Jiang, Zhixin Shu +3Video World Models

  3. Astronex-World 1.0: Real-Time Interactive World Model Foundation

    Sep 17, 2026Xin Zhou, Cong MiaoVideo World ModelsAction Space