Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, current approaches still struggle to generate physically plausible object interactions and lack object-level control mechanisms. To address these limitations, we introduce KineMask, an approach for video generation that enables realistic rigid body control, interactions, and effects. Given a single image and a specified object velocity, our method generates videos with inferred motions and future object interactions. We propose a two-stage training strategy that gradually removes future motion supervision via object masks. Using this strategy we train video diffusion models (VDMs) on synthetic scenes of simple interactions and demonstrate significant improvements and generalization to rigid body and hand-object interactions in real scenes. Furthermore, KineMask integrates low-level motion control with high-level textual conditioning via predicted scene descriptions, leading to support for synthesis of complex dynamical phenomena. Our experiments show that KineMask generalizes to different VDMs and achieves strong improvements over recent models of comparable size. Ablation studies further highlight the complementary roles of low- and high-level conditioning in VDMs.
Figures & tables
Figure 1 : We enable object-based control with a novel training strategy. Paired with synthetic data constructed for the task, KineMask enables pretrained VDMs to synthesize realistic rigid body interactions in real-world input scenes.
Figure 2 : KineMask pipeline. We encode our low-level control signal as a mask encoding the velocity of the moving objects, to train a ControlNet (left) in two stages using Blender-generated videos of objects in motion. In the first stage, we train with all mask frames as control. In the second stage, we randomly drop parts of the final mask frames. Additionally, we provide a high-level textual control extracted by a VLM. At inference (right), we construct the low-level conditioning with SAM and use GPT to infer high-level outcomes of object motion from a single frame.
Figure 3 : Comparison with CogVideoX. While CogVideoX often suffers from several failure modes, such as hallucinations and incorrect motions, KineMask follows target motion and generates realistic object interactions. We improve object interactions in collisions and show causal effects of object motion.
Figure 4 : Degrees of freedom. We show control of different aspects of KineMask outputs. We can choose different directions (left), speed (middle), and objects to move (right), opening potential for world modeling.
Figure 5 : We widely outperform baselines on motion fidelity, interaction quality, and overall physical consistency.
Figure 6 : Impact of training data. While KineMask trained on Simple Motion is able to generalize to Real World images, the lack of object interactions in Simple Motion results in hallucinations (top). Training on Interactions results in collisions and plausible motion of pushed objects (bottom).
Simple Motion
Method
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
CogVideoX
158.3
601.1
1504.6
0.051
2nd stage only
86.3
288.8
201.0
0.237
Ours (1st + 2nd stage)
47.2
160.3
199.8
0.367
1st stage w/ masks
24.9
89.7
165.6
0.684
Table 1 : Ablation on training . We show that our two-stage training strategy considerably boosts results and approaches an upper bound exploiting privileged information (1st stage w/ masks).
Figure 7 : On top, we show how different velocities impact motion effects, suggesting causal understanding. At the bottom, we show that the environment impacts results, for the same velocity but different objects, suggesting mass awareness.
Interactions
Method
Training
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
CogVideoX
-
344.6
807.3
3514.9
0.192
CogVideoX
Interactions
201.9
368.2
162.9
0.243
KineMask
Simple Motion
166.2
301.1
160.5
0.334
KineMask
Interactions
158.7
250.7
143.8
0.355
Table 2 : Effects of data. Compared with models trained on Simple Motion , training on Interactions data considerably boosts performance in all metrics.
Figure 8 : Multi-Object Control. Generation examples of KineMask for multiple object control and interactions, using velocity masks for different objects as input generates objects colliding with each other or with other objects in the scene.
Figure 9 : Generalization. KineMask trained on Interactions is also able to generalize its knowledge to Hand-Tool-Object (Left) and Hand-Object (Right) interactions, showing plausible generations that respect causality.
Figure 10 : Sequential Generation. Generation examples applying KineMask sequentially, by combining our velocity masks and text only.
Figure 11 : Impact of text. Training with rich captions c makes KineMask generate effects that go beyond the synthetic data used for training, exploiting the prior knowledge of the VDM. While we prompt both models with cinfer , KineMask trained with c (bottom row) is able to generate complex effects of the interactions, such as crashing objects, and water effects. Full prompts in Appendix.
Interactions
Train
Infer
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
c∅
c∅
158.7
250.7
143.8
0.355
c∅
cinfer
174.4
238.8
161.3
0.356
c
cinfer
160.9
231.3
174.4
0.376
Table 3: Ablation on prompts. Training with c improve object consistency and quality on synthetic data.
Figure 12 : All versions of KineMask outperform the original models, on motion fidelity, interaction quality and overall physical consistency. Showing its applicability, and generalization, bringing improvements to different video diffusion models.
Figure 13 : Qualitative comparison with Wan and Cosmos. KineMask-Wan and KineMask-Cosmos follows target motion and generates realistic object interactions, while the original backbones show implausible motions. This demonstrates that apart from our original implementation on top of CogVideoX, KineMask generalize to different VDM’s bringing improvements to all models.
Figure 14 : In-context learning examples. Providing examples to GPT helps the generation of the text cinfer following a desired format.
Figure 15 : User study interface. We collect pairwise user preference on direction, interactions realism, and consistency of objects.
Figure 16 : Qualitative comparison with all baselines.
Figure 17 : Qualitative comparison of KineMask vs PhysGen3D.
Interactions
Backbone
Setup
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
Wan
Baseline
431.4
977.7
3649.8
0.192
KineMask
211.0
475.6
968.4
0.257
CogVideoX
Baseline
344.6
807.3
3514.9
0.192
KineMask
158.7
250.7
143.8
0.355
Cosmos
Baseline
1463.6
1287.3
1285.7
0.203
Table 4 : Comparison of different model backbones. We compare all versions of KineMask using different base models.
Figure 18 : Additional Comparisons on Real-World Videos. Real and Generation examples for our own data (left) and Physics-IQ (right)
Physics IQ
Our Data
Method
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
Force-Prompting
27.76
652.5
860.4
0.358
8029.4
469.7
1833.9
0.217
TORA
28.05
667.0
1327.6
0.330
9273.1
465.2
4107.7
0.193
CogvideoX (zero-shot)
88.87
581.3
1608.1
0.282
11367.5
609.4
3404.4
0.227
CogvideoX (finetuned)
65.86
357.1
331.6
0.220
6032.2
319.22
405.62
0.201
KineMask
39.67
212.1
431.81
0.428
5551.9
295.7
422.38
0.339
Table 5 : Additional Quantitative Results. We report additional comparisons for our main model Kinemask-CogVideoX against the main baselines and the zero-shot and finetuned backbone model CogvideoX.
Table 6 : Additional ablation studies. We report an additional encoding strategies for our conditioning masks in Table ( 6(a) ), proving that our design choice of using vt is best. Furthermore, we showcase that removing our non-uniform sampling on interactions-rich frames during dropout harms performance ( 6(b) ).
Figure 19 : Role of Text. Generation examples when the mask and prompt disagree. Left, the prompt describes a collision while the velocity is not enough to create the interaction. Right, the prompt does not describe an interaction but the velocities of both bottles are enough to create a collision.
Figure 20 : Motion control. Our mask-based control is robust to ambiguous scenes, while others suffer from hallucinations due to the ambiguous control signal.
Figure 21 : Failure Cases. We noticed that KineMask fails to generate collisions with objects that do not have a considerable height and in ambiguous setups.
Figure 22 : Additional qualitative examples of KineMask.