While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To address this, we introduce a framework to generate a scalable synthetic dataset of controllable interactions. Our pipeline leverages a structured taxonomy and state-of-the-art image editing models to create explicit start' and end' state images, which serve as visual anchors for the interaction. To generate a seamless video utilizing these anchors, we propose State-Guided Sampling (SGS), a novel sampling technique that mitigates artifacts common in naive conditional generation. Furthermore, we develop and validate a new automated evaluation system that aligns with human judgments to ensure data quality. Experiments show that fine-tuning a base model on our dataset significantly enhances its ability to generate plausible interactions.
Figures & tables
Figure 1: Overview of Interaction-Centric Video Dataset Generation Pipeline.
Figure 2: Qualitative Evaluation. The red arrow indicates abrupt frame changes, while the plus ( + ) denotes our fine-tuned model.
Figure 3: VLM-Human Rating Correlation.
F1 (Good)
F1 (Bad)
Macro F1
AUC
VideoPhy2
0.00
0.80
0.40
0.57
Base
0.53
0.80
0.66
0.71
(a) + Prefilter
0.59
0.77
0.68
0.72
(b) + PC score
0.60
0.78
0.69
0.72
(c) + Surprise
0.61
0.79
0.70
0.73
(d) + PP
0.61
0.79
0.70
0.76
Table 1: Ablation Study on SVM Inputs. PP and QC denote the features from Plausibility Probe and Quality Classifier, respectively.
Figure 4: Effectiveness of State-Guided Sampling. The proposed SGS effectively resolves the unnatural scene transitions.
Model
VLM-Assisted Score ( ↑ )
Temporal Artifact ( ↓ )
VideoPhy2 ( ↑ )
Clarity
Causality
Plausibility
Continuity
Average
SA
PC
PhyI2V 1
2.52
2.34
2.44
3.40
2.68
0.44
0.26
0.58
HunyuanVideo
1.71
1.50
3.06
4.16
2.61
0.09
0.25
0.77
FLF
3.10
3.06
2.87
3.20
3.06
0.68
0.31
0.56
Wan 2.1
I2V
2.90
2.86
2.98
3.87
3.15
0.13
0.30
0.56
SGS
3.00
2.94
2.91
3.96
3.20
0.12
0.26
0.53
Table 2: Quantitative Evaluation. SA and PC denote Semantic Adherence and Physical Commonsense, respectively. The bold represents the best, and the underline does the second best.
Vbench
Human Comparison
Motion Smoothness
Dynamic Degree
Aesthetic Quality
Rank ( ↓ )
WR
PhyI2V
0.99
0.57
0.57
2.90
8%
Hunyuan
0.99
0.20
0.63
3.43
2%
Wan2.1 I2V
0.98
0.64
0.63
1.98
29%
+Fine-tune
0.98
0.73
0.63
1.68
53%
Table 3: VBench (Motion) and Human Evaluation.
Clar.
Caus.
Plau.
Cont.
Rank( ↓ )
Good
Temporal Artifact( ↓ )
FLF
2.96
2.64
2.82
2.41
4.08
0.15
0.66
I2V
2.64
2.36
3.01
4.22
3.71
0.22
0.13
α=0.3
3.04
2.72
3.05
4.21
3.05
0.30
0.26
α=0.4
3.13
2.81
3.11
4.30
2.98
0.35
0.20
α=0.5
3.14
2.86
3.07
4.56
3.05
0.38
0.11
α=0.6
2.95
2.65
3.08
4.55
3.38
0.31
0.11
Table 4: Ablation Study on the Initial I2V Weight, α , for SGS.
PhyGenBench
Human Comparison
Mechanics
Optics
Thermal
Material
Average
Rank ( ↓ )
WR
PhyI2V
0.51
0.62
0.54
0.49
0.55
2.79
19%
Hunyuan
0.43
0.57
0.42
0.28
0.44
2.93
12%
Wan2.1 I2V
0.49
0.59
0.54
0.39
0.51
2.38
27%
+Fine-tune
0.47
0.63
0.57
0.41
0.52
1.90
42%
Table 5: Evaluation on the PhyGenBench. (I2V)
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Precision
Recall
F1-score
Support
Bad
0.82
0.78
0.80
530
Good
0.61
0.65
0.63
268
Macro Avg
0.71
0.72
0.71
798
Weighted Avg
0.75
0.74
0.74
798
Accuracy
0.74
ROC AUC
0.77
Appendix
Table 6: Overall Our SVM Evaluator Performance.
Figure 5: The hierarchical structure of our Actor-Object and Interactions-Actions Taxonomy , forming the basis for generating diverse interaction scenarios.
Figure 6: An example of our data generation pipeline. From a sampled interaction (Extinguish_fire) and objects (Child, Gum), our method completes the tuple, generates detailed prompts for the ‘start’ and ‘end’ states, synthesizes the corresponding images, and finally generates the video representing the state transition.
Figure 7: Failure cases produced by naive sampling methods. These examples illustrate the typical artifacts that our State-Guided Sampling (SGS) is designed to resolve. (a) The I2V model fails to depict the state change, resulting in a static and visually inconsistent video. (b) The FLF model creates an abrupt and unnatural transition with noticeable visual artifacts.
Correlation
Accuracy
Precision
Recall
F1-score
0.637
0.865
0.572
0.820
0.674
Appendix
Table 7: Performance of the temporal artifact detector. The Pearson Correlation is calculated against the raw (1-5) human-rated Continuity scores. The binary classification metrics are based on a threshold where a human Continuity score ≤2.0 defines the positive class.
Figure 8: Qualitative examples of the ‘ghosting effect’. Naive Constant Interpolation (left column) results in an unnatural, translucent overlay in the final frames. In contrast, our proposed SGS (right column) resolves this artifact by dynamically adjusting model influence, producing a clear and temporally coherent final state.
Figure 9: Qualitative Baseline Comparison.
Figure 10: Qualitative Baseline Comparison.
Figure 11: Qualitative Baseline Comparison.
Figure 12: Qualitative comparison of SGS against naive sampling methods.
Figure 13: Qualitative comparison on PhyGenBench. While our model produces plausible videos in (a-c), only PhyI2V receives high scores, suggesting a scoring bias where its VLM-based refinement games the VLM-based evaluator. In contrast, PhyI2V’s high score in (d) stems from a valid advantage in world knowledge (litmus solution chemistry).
Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact, an intended action is skipped, a placed object keeps drifting, or a support relation breaks. Our starting point is that standard frame-first denoising updates every latent region at every step, even when the prompt implies that only a local interaction should be active. We introduce Event-Driven Video Generation (EVD), a small DiT-compatible intervention that gives the sampler an explicit event signal. A lightweight head predicts token-level event activity; training losses tie that activity to latent state change; and event-gated sampling, with hysteresis and an early-step schedule, applies the update field mainly where an interaction is forming. On EVD-Bench, EVD improves human preference and VBench dynamics for state persistence, spatial accuracy, support relations, and contact stability, while keeping appearance quality comparable to the base model. The results suggest that a modest amount of event structure can correct several interaction failures that otherwise remain hidden behind good frame-level appearance.
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.
Vedant Shah, Onkar Susladkar, Tushar Prakash +5
University of Illinois Urbana-Champaign · Sony Research India
Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increasingly demand multi-person videos with rich social interactions, including conversations, gestures, and coordinated actions. However, existing models offer no explicit control over interactions, such as who performs which action, when it occurs, and toward whom it is directed. This often results in wrong person performing unintended actions (actor-action mismatch), disordered social dynamics, and wrong action targets. To address these challenges, we present SocialDirector, a training-free interaction controller that enhances the generation model by modulating cross-attention maps. SocialDirector contains two modules: Social Actor Masking and Directional Reweighting. Social Actor Masking constrains each person's visual tokens to attend only to their own textual descriptions via a spatiotemporal mask, avoiding actor-action mismatch and disordered social dynamics. Directional Reweighting amplifies attention to directional words (e.g., "leftward", "right"), leading each action towards its intended target. To evaluate generated social interactions, we annotate existing datasets with interaction descriptions and build a fully automated evaluation pipeline powered by open-source VLMs. Experiments on different video generation models show that SocialDirector significantly improves interaction fidelity and approaches the upper bound set by real videos.
Liangyang Ouyang, Ruicong Liu, Caixin Kang +2
1The University of Tokyo · 2Shanda AI Research Tokyo