Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation. Project website: https://glab-caltech.github.io/dynatokens/
Figures & tables
Figure 1 : DynaTokens enables camera-controlled video models to learn dynamics via test-time training. (a) DynaTokens introduces learnable tokens that are cross-attended by the video patch tokens in every transformer block. (b) For each scene, we train DynaTokens on a small set of camera trajectories ( ≈ 15), enabling inference on novel, unseen trajectories. (c) Visualization of the attention weights for a dynamic token, corroborating the localized dynamics hypothesis.
Figure 2 : Motivating example.
Figure 3 : Top: Comparison of DynaTokens with state-of-the-art camera-controlled video models. DynaTokens consistently enables dynamics, whereas existing models struggle, either keeping the object static, show incorrect motion, or produce low-quality videos. Bottom: Comparison across different test-time training strategies. Only DynaTokens is able to correctly learn dynamics while maintaining camera control. The camera control is overlayed on the frames using keyboard schema, “wasd” means translation and arrows mean rotation.
Table 1 : (a) Comparing DynaTokens with state-of-the-art camera-controlled video models. We evaluate Dynamic Spatial Relation (DSR) and Motion Order Understanding (MOU) on VBench2 and Motion Accuracy (MA) on WorldScore, as well as camera. Current state-of-the-art models struggle greatly with dynamics. DynaTokens, via test-time training, effectively learns dynamics while maintaining camera. (b) DynaTokens outperforms alternative test-time training methods, LoRA, finetuning, and TTT layer for dynamics. Standard errors are in Table 4 and Table 5 in Appendix.
Figure 5
Figure 4 : Inference Time
Trajectories
Dynamics
Camera
Seen
0.83
0.85
Unseen
0.81
0.86
Table 2 : DynaTokens demonstrates strong generalization, evaluated on VBench2.
Figure 5 : DynaTokens enables diverse dynamics, including physics and stylistic changes, whereas current state-of-the-art camera-controlled video models like Lingbot and Genie 3 struggle.
DSR
MOU
Cam
DynaTokens
1.00
0.72
0.88
Single token
0.57
0.56
0.89
Reduce token dim (32)
0.65
0.72
0.88
Increase token dim (128)
0.60
0.72
0.87
No context noising
0.67
0.83
0.84
Additive
0.91
0.67
0.85
Table 3 : Ablations on dynamic token design evaluated on VBench2. The current design yields best holistic performance on dynamics and camera. DSR: Dynamic Spatial Relation. MOU: Motion Order Understanding.
Figure 10
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
VBench2
WorldScore
DSR
MOU
Cam
MA
Cam
Lyra2
0.54 (0.071)
0.17 (0.042)
0.87 (0.017)
1.11 (0.226)
0.70 (0.015)
Lingbot
0.68 (0.071)
0.25 (0.058)
0.82 (0.017)
4.37 (0.605)
0.64 (0.023)
WorldPlay
0.56 (0.070)
0.07 (0.032)
0.88 (0.009)
1.95 (0.273)
0.69 (0.013)
Hydra
0.51 (0.066)
0.18 (0.050)
0.61 (0.019)
2.51 (0.520)
0.56 (0.029)
LiveWorld
0.44 (0.061)
0.35 (0.062)
0.82 (0.016)
4.49 (0.817)
0.70 (0.013)
Appendix
Table 4 : Comparison of DynaTokens with camera-controlled video models on VBench2 and WorldScore with standard errors across test samples. DynaTokens’ advantage is significant.
VBench2
WorldScore
DSR
MOU
Cam
MA
Cam
Dynatoken
1.00 (0.000)
0.72 (0.082)
0.88 (0.011)
4.46 (0.538)
0.67 (0.060)
LoRA
0.61 (0.113)
0.22 (0.101)
0.83 (0.008)
2.79 (0.523)
0.56 (0.062)
LoRA-no PRoPE
0.41 (0.123)
0.06 (0.056)
0.84 (0.008)
3.62 (0.337)
0.67 (0.056)
Finetuning
0.59 (0.123)
0.67 (0.114)
0.83 (0.021)
2.66 (0.358)
0.54 (0.064)
TTT Layer
0.41(0.110)
0.11 (0.076)
0.88 (0.007)
3.79 (0.384)
0.66 (0.061)
Appendix
Table 5 : Comparison of DynaTokens with other test-time training methods, including LoRA, block finetuning, and TTT layer, with standard errors across evaluation examples. DynaTokens’ advantage is significant.
VBench
WorldScore
DSR
MOU
Cam
MA
Cam
DynaTokens
1.00
0.75
0.86
5.15
0.69
VPT/APT
0.50
0.17
0.84
2.35
0.67
Appendix
Table 6 : Comparison of DynaTokens and VPT/APT. Naive prefix tuning methods cannot disentangle dynamics from camera control, and struggles to learn dynamics effectively.
IoU
IoU-d1
DynaTokens
0.342
0.511
LoRA
0.006
0.017
Text
0.011
0.024
Chance
0.010
0.022
Appendix
Table 7 : IoU and IoU-d1 (relaxed to count all distnace ≤ 1 neighbors as correct) of DynaTokens attention map, attention map delta before and after LoRA, the attention map from text keywords corresponding to the moving object (such as the word “dog”), as well as a random attention mask pattern. DynaTokens has a significantly higher overlap with the SAM mask of the moving object, while LoRA and text perform similarly to random chance.
# trajectories
VBench
WorldScore
DSR
MOU
Cam
MA
Cam
3
0.67
0.50
0.80
4.02
0.67
6
0.75
0.67
0.80
4.68
0.68
12
1.00
0.75
0.86
4.81
0.68
15
1.00
0.75
0.87
5.15
0.68
Appendix
Table 8 : Studying the effect of the number of training trajectories. Even when training on as few as 3 trajectories, DynaTokens shows a considerable advantage on MOU (0.50 vs. 0.35 for the best baseline, LiveWorld).
# artifact swap
VBench
WorldScore
DSR
MOU
Cam
MA
Cam
0 (original)
1.00
0.75
0.87
5.15
0.68
1 (7%)
1.00
0.75
0.86
4.99
0.66
5 (33%)
1.00
0.67
0.86
4.50
0.66
Appendix
Table 9 : Studying the effect of the quality of training trajectories by swapping k good trajectories with k trajectories with artifacts. Even with 33% trajectories swapped to corrupted, DynaTokens’ performance remains high.
Figure 6 : Examples of curated training examples. Curating one sample takes 20 GPU minutes (A100) and cost $1.52 in API calls (Kling and Gemini). The curated data contains correct dynamics but may include background inconsistencies. Our test-time training approach resolves such inconsistencies by leveraging model priors.
Figure 7 : Additional qualitative examples of comparison with state-of-the-art camera-controlled video models.
Figure 8 : Additional qualitative examples of comparison with state-of-the-art camera-controlled video models.
Figure 9 : Additional qualitative examples of comparison with state-of-the-art camera-controlled video models.
Figure 10 : Additional qualitative examples of comparison with state-of-the-art camera-controlled video models.
Figure 11 : Additional qualitative examples of comparison across test-time methods.
Figure 12 : Additional qualitative examples of comparison across test-time methods.
Figure 13 : Additional qualitative examples for physical dynamics.
Figure 14 : DynaTokens can enable dynamics of multiple objects.
Figure 15 : Attention map visualization of DynaTokens, LoRA, and the text keyword of the moving object (dog).
Figure 16 : Attention map visualization of DynaTokens, LoRA, and the text keyword of the moving object (kangaroo).
Figure 17 : Attention map visualization of DynaTokens, LoRA, and the text keyword of the moving object (person).
Figure 18 : Attention map visualization of DynaTokens for a complex scene with four moving objects (four ducks moving in different directions).