Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, current approaches still struggle to generate physically plausible object interactions and lack object-level control mechanisms. To address these limitations, we introduce KineMask, an approach for video generation that enables realistic rigid body control, interactions, and effects. Given a single image and a specified object velocity, our method generates videos with inferred motions and future object interactions. We propose a two-stage training strategy that gradually removes future motion supervision via object masks. Using this strategy we train video diffusion models (VDMs) on synthetic scenes of simple interactions and demonstrate significant improvements and generalization to rigid body and hand-object interactions in real scenes. Furthermore, KineMask integrates low-level motion control with high-level textual conditioning via predicted scene descriptions, leading to support for synthesis of complex dynamical phenomena. Our experiments show that KineMask generalizes to different VDMs and achieves strong improvements over recent models of comparable size. Ablation studies further highlight the complementary roles of low- and high-level conditioning in VDMs.
Figures & tables
Figure 1 : We enable object-based control with a novel training strategy. Paired with synthetic data constructed for the task, KineMask enables pretrained VDMs to synthesize realistic rigid body interactions in real-world input scenes.
Figure 2 : KineMask pipeline. We encode our low-level control signal as a mask encoding the velocity of the moving objects, to train a ControlNet (left) in two stages using Blender-generated videos of objects in motion. In the first stage, we train with all mask frames as control. In the second stage, we randomly drop parts of the final mask frames. Additionally, we provide a high-level textual control extracted by a VLM. At inference (right), we construct the low-level conditioning with SAM and use GPT to infer high-level outcomes of object motion from a single frame.
Figure 3 : Comparison with CogVideoX. While CogVideoX often suffers from several failure modes, such as hallucinations and incorrect motions, KineMask follows target motion and generates realistic object interactions. We improve object interactions in collisions and show causal effects of object motion.
Figure 4 : Degrees of freedom. We show control of different aspects of KineMask outputs. We can choose different directions (left), speed (middle), and objects to move (right), opening potential for world modeling.
Figure 5 : We widely outperform baselines on motion fidelity, interaction quality, and overall physical consistency.
Figure 6 : Impact of training data. While KineMask trained on Simple Motion is able to generalize to Real World images, the lack of object interactions in Simple Motion results in hallucinations (top). Training on Interactions results in collisions and plausible motion of pushed objects (bottom).
Simple Motion
Method
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
CogVideoX
158.3
601.1
1504.6
0.051
2nd stage only
86.3
288.8
201.0
0.237
Ours (1st + 2nd stage)
47.2
160.3
199.8
0.367
1st stage w/ masks
24.9
89.7
165.6
0.684
Table 1 : Ablation on training . We show that our two-stage training strategy considerably boosts results and approaches an upper bound exploiting privileged information (1st stage w/ masks).
Figure 7 : On top, we show how different velocities impact motion effects, suggesting causal understanding. At the bottom, we show that the environment impacts results, for the same velocity but different objects, suggesting mass awareness.
Interactions
Method
Training
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
CogVideoX
-
344.6
807.3
3514.9
0.192
CogVideoX
Interactions
201.9
368.2
162.9
0.243
KineMask
Simple Motion
166.2
301.1
160.5
0.334
KineMask
Interactions
158.7
250.7
143.8
0.355
Table 2 : Effects of data. Compared with models trained on Simple Motion , training on Interactions data considerably boosts performance in all metrics.
Figure 8 : Multi-Object Control. Generation examples of KineMask for multiple object control and interactions, using velocity masks for different objects as input generates objects colliding with each other or with other objects in the scene.
Figure 9 : Generalization. KineMask trained on Interactions is also able to generalize its knowledge to Hand-Tool-Object (Left) and Hand-Object (Right) interactions, showing plausible generations that respect causality.
Figure 10 : Sequential Generation. Generation examples applying KineMask sequentially, by combining our velocity masks and text only.
Figure 11 : Impact of text. Training with rich captions c makes KineMask generate effects that go beyond the synthetic data used for training, exploiting the prior knowledge of the VDM. While we prompt both models with cinfer , KineMask trained with c (bottom row) is able to generate complex effects of the interactions, such as crashing objects, and water effects. Full prompts in Appendix.
Interactions
Train
Infer
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
c∅
c∅
158.7
250.7
143.8
0.355
c∅
cinfer
174.4
238.8
161.3
0.356
c
cinfer
160.9
231.3
174.4
0.376
Table 3: Ablation on prompts. Training with c improve object consistency and quality on synthetic data.
Figure 12 : All versions of KineMask outperform the original models, on motion fidelity, interaction quality and overall physical consistency. Showing its applicability, and generalization, bringing improvements to different video diffusion models.
Figure 13 : Qualitative comparison with Wan and Cosmos. KineMask-Wan and KineMask-Cosmos follows target motion and generates realistic object interactions, while the original backbones show implausible motions. This demonstrates that apart from our original implementation on top of CogVideoX, KineMask generalize to different VDM’s bringing improvements to all models.
Figure 14 : In-context learning examples. Providing examples to GPT helps the generation of the text cinfer following a desired format.
Figure 15 : User study interface. We collect pairwise user preference on direction, interactions realism, and consistency of objects.
Figure 16 : Qualitative comparison with all baselines.
Figure 17 : Qualitative comparison of KineMask vs PhysGen3D.
Interactions
Backbone
Setup
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
Wan
Baseline
431.4
977.7
3649.8
0.192
KineMask
211.0
475.6
968.4
0.257
CogVideoX
Baseline
344.6
807.3
3514.9
0.192
KineMask
158.7
250.7
143.8
0.355
Cosmos
Baseline
1463.6
1287.3
1285.7
0.203
Table 4 : Comparison of different model backbones. We compare all versions of KineMask using different base models.
Figure 18 : Additional Comparisons on Real-World Videos. Real and Generation examples for our own data (left) and Physics-IQ (right)
Physics IQ
Our Data
Method
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
MSE ↓
FVD ↓
FMVD ↓
IoU ↑
Force-Prompting
27.76
652.5
860.4
0.358
8029.4
469.7
1833.9
0.217
TORA
28.05
667.0
1327.6
0.330
9273.1
465.2
4107.7
0.193
CogvideoX (zero-shot)
88.87
581.3
1608.1
0.282
11367.5
609.4
3404.4
0.227
CogvideoX (finetuned)
65.86
357.1
331.6
0.220
6032.2
319.22
405.62
0.201
KineMask
39.67
212.1
431.81
0.428
5551.9
295.7
422.38
0.339
Table 5 : Additional Quantitative Results. We report additional comparisons for our main model Kinemask-CogVideoX against the main baselines and the zero-shot and finetuned backbone model CogvideoX.
Table 6 : Additional ablation studies. We report an additional encoding strategies for our conditioning masks in Table ( 6(a) ), proving that our design choice of using vt is best. Furthermore, we showcase that removing our non-uniform sampling on interactions-rich frames during dropout harms performance ( 6(b) ).
Figure 19 : Role of Text. Generation examples when the mask and prompt disagree. Left, the prompt describes a collision while the velocity is not enough to create the interaction. Right, the prompt does not describe an interaction but the velocities of both bottles are enough to create a collision.
Figure 20 : Motion control. Our mask-based control is robust to ambiguous scenes, while others suffer from hallucinations due to the ambiguous control signal.
Figure 21 : Failure Cases. We noticed that KineMask fails to generate collisions with objects that do not have a considerable height and in ambiguous setups.
Figure 22 : Additional qualitative examples of KineMask.
We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence conditioned on an input image. CoMoGen introduces a lightweight MaskAdapter that encodes binary mask sequences into a latent residual signal, injected into the Multi Modal Diffusion Transformer (MMDiT) model through a cosine-weighted schedule. Unlike the hierarchical coarse-to-fine design of UNet architectures, MMDiT operates as a sequence of uniform transformer blocks, making it difficult to identify which layers are responsible for the motion generation. Therefore, we propose a novel way to determine "Motion Layers" operating in the attention space of MMDiT. We fine-tune the model by using Low-Rank Adaptation (LoRA) to the Motion Layers, without requiring any architecture change in the MMDiT. This selective adaptation enables our method to focus on motion-critical components, yielding reduced computational cost. Despite its simplicity, CoMoGen enables precise subject motion and plausible interactions with surrounding humans, objects, and scenes. Comprehensive experiments on different datasets show that CoMoGen consistently outperforms prior controllable video generation methods and achieves state-of-the-art performance in motion fidelity and perceptual realism. Project page: mericadil.github.io/CoMoGen.
Adil Meric, Lin Geng Foo, Mert Kiray +3
Technical University of Munich · Max Planck Institute for Informatics, Saarland Informatics Campus · Munich Center for Machine Learning (MCML) +1
Video-based embodied world models provide an appealing substrate for robotic manipulation by predicting future states, yet current approaches remain limited by a fundamental entanglement: accurately modeling dynamics typically requires low-level temporal reasoning, while producing high-resolution frames demands expansive visual synthesis according to high-level semantics. This entanglement results in slow inference speed for iterative planning or too coarse predictions to retain contact-rich details. To solve this dilemma, we present Disentangled Video Generation World Model (DVG-WM), an efficient framework that explicitly decomposes world modeling into dynamics learning and visual synthesis. Conditioned on an initial observation and a language instruction, our model first generates a plausible sequence of intermediate visual states to preview the physical interaction and refines them to obtain high-fidelity videos. Furthermore, an efficient cascading mechanism is proposed, where DVG-WM uses flow matching to directly map the dynamics to video latents, and introduces a latent degradation mechanism to regenerate contact-rich details. Experiments on LIBERO and real-world platforms demonstrate improved video quality with up to 3.97 times acceleration, validating that disentangled video generation can be an efficient embodied world model for robotic manipulation.
Ziyu Shan, Zhenyu Wu, Xiaofeng Wang +2
Nanyang Technological University, Singapore · Beijing University of Posts and Telecommunications, Beijing, China · GigaAI
Video generation models achieve high visual quality but often struggle to generate physics-aware videos. Unlike rigid-body motion, which can be described by explicit trajectories or formulas, complex deformation dynamics remain challenging to synthesize. We observe that a lack of physical reasoning for localizing dynamic areas allows irrelevant regions to dilute the model's attention, leading to generation failure. In this paper, we propose DeforM, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions. To reason about and localize these critical regions, we introduce a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks. For physical guidance, we develop two alternative strategies: DeforM-Free for training-free mechanism analysis and DeforM-Injection as a powerful training-based generator. Experimental results demonstrate that DeforM improves the realism of generated deformation scenarios, outperforming baseline models in both visual quality and physical consistency.
Yunyi Li, Yu Qiao, Yaohui Wang +1
Shanghai Artificial Intelligence Laboratory, Shanghai, China