While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To address this, we introduce a framework to generate a scalable synthetic dataset of controllable interactions. Our pipeline leverages a structured taxonomy and state-of-the-art image editing models to create explicit start' and end' state images, which serve as visual anchors for the interaction. To generate a seamless video utilizing these anchors, we propose State-Guided Sampling (SGS), a novel sampling technique that mitigates artifacts common in naive conditional generation. Furthermore, we develop and validate a new automated evaluation system that aligns with human judgments to ensure data quality. Experiments show that fine-tuning a base model on our dataset significantly enhances its ability to generate plausible interactions.
Figures & tables
Figure 1: Overview of Interaction-Centric Video Dataset Generation Pipeline.
Figure 2: Qualitative Evaluation. The red arrow indicates abrupt frame changes, while the plus ( + ) denotes our fine-tuned model.
Figure 3: VLM-Human Rating Correlation.
F1 (Good)
F1 (Bad)
Macro F1
AUC
VideoPhy2
0.00
0.80
0.40
0.57
Base
0.53
0.80
0.66
0.71
(a) + Prefilter
0.59
0.77
0.68
0.72
(b) + PC score
0.60
0.78
0.69
0.72
(c) + Surprise
0.61
0.79
0.70
0.73
(d) + PP
0.61
0.79
0.70
0.76
Table 1: Ablation Study on SVM Inputs. PP and QC denote the features from Plausibility Probe and Quality Classifier, respectively.
Figure 4: Effectiveness of State-Guided Sampling. The proposed SGS effectively resolves the unnatural scene transitions.
Model
VLM-Assisted Score ( ↑ )
Temporal Artifact ( ↓ )
VideoPhy2 ( ↑ )
Clarity
Causality
Plausibility
Continuity
Average
SA
PC
PhyI2V 1
2.52
2.34
2.44
3.40
2.68
0.44
0.26
0.58
HunyuanVideo
1.71
1.50
3.06
4.16
2.61
0.09
0.25
0.77
FLF
3.10
3.06
2.87
3.20
3.06
0.68
0.31
0.56
Wan 2.1
I2V
2.90
2.86
2.98
3.87
3.15
0.13
0.30
0.56
SGS
3.00
2.94
2.91
3.96
3.20
0.12
0.26
0.53
Table 2: Quantitative Evaluation. SA and PC denote Semantic Adherence and Physical Commonsense, respectively. The bold represents the best, and the underline does the second best.
Vbench
Human Comparison
Motion Smoothness
Dynamic Degree
Aesthetic Quality
Rank ( ↓ )
WR
PhyI2V
0.99
0.57
0.57
2.90
8%
Hunyuan
0.99
0.20
0.63
3.43
2%
Wan2.1 I2V
0.98
0.64
0.63
1.98
29%
+Fine-tune
0.98
0.73
0.63
1.68
53%
Table 3: VBench (Motion) and Human Evaluation.
Clar.
Caus.
Plau.
Cont.
Rank( ↓ )
Good
Temporal Artifact( ↓ )
FLF
2.96
2.64
2.82
2.41
4.08
0.15
0.66
I2V
2.64
2.36
3.01
4.22
3.71
0.22
0.13
α=0.3
3.04
2.72
3.05
4.21
3.05
0.30
0.26
α=0.4
3.13
2.81
3.11
4.30
2.98
0.35
0.20
α=0.5
3.14
2.86
3.07
4.56
3.05
0.38
0.11
α=0.6
2.95
2.65
3.08
4.55
3.38
0.31
0.11
Table 4: Ablation Study on the Initial I2V Weight, α , for SGS.
PhyGenBench
Human Comparison
Mechanics
Optics
Thermal
Material
Average
Rank ( ↓ )
WR
PhyI2V
0.51
0.62
0.54
0.49
0.55
2.79
19%
Hunyuan
0.43
0.57
0.42
0.28
0.44
2.93
12%
Wan2.1 I2V
0.49
0.59
0.54
0.39
0.51
2.38
27%
+Fine-tune
0.47
0.63
0.57
0.41
0.52
1.90
42%
Table 5: Evaluation on the PhyGenBench. (I2V)
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Precision
Recall
F1-score
Support
Bad
0.82
0.78
0.80
530
Good
0.61
0.65
0.63
268
Macro Avg
0.71
0.72
0.71
798
Weighted Avg
0.75
0.74
0.74
798
Accuracy
0.74
ROC AUC
0.77
Appendix
Table 6: Overall Our SVM Evaluator Performance.
Figure 5: The hierarchical structure of our Actor-Object and Interactions-Actions Taxonomy , forming the basis for generating diverse interaction scenarios.
Figure 6: An example of our data generation pipeline. From a sampled interaction (Extinguish_fire) and objects (Child, Gum), our method completes the tuple, generates detailed prompts for the ‘start’ and ‘end’ states, synthesizes the corresponding images, and finally generates the video representing the state transition.
Figure 7: Failure cases produced by naive sampling methods. These examples illustrate the typical artifacts that our State-Guided Sampling (SGS) is designed to resolve. (a) The I2V model fails to depict the state change, resulting in a static and visually inconsistent video. (b) The FLF model creates an abrupt and unnatural transition with noticeable visual artifacts.
Correlation
Accuracy
Precision
Recall
F1-score
0.637
0.865
0.572
0.820
0.674
Appendix
Table 7: Performance of the temporal artifact detector. The Pearson Correlation is calculated against the raw (1-5) human-rated Continuity scores. The binary classification metrics are based on a threshold where a human Continuity score ≤2.0 defines the positive class.
Figure 8: Qualitative examples of the ‘ghosting effect’. Naive Constant Interpolation (left column) results in an unnatural, translucent overlay in the final frames. In contrast, our proposed SGS (right column) resolves this artifact by dynamically adjusting model influence, producing a clear and temporally coherent final state.
Figure 9: Qualitative Baseline Comparison.
Figure 10: Qualitative Baseline Comparison.
Figure 11: Qualitative Baseline Comparison.
Figure 12: Qualitative comparison of SGS against naive sampling methods.
Figure 13: Qualitative comparison on PhyGenBench. While our model produces plausible videos in (a-c), only PhyI2V receives high scores, suggesting a scoring bias where its VLM-based refinement games the VLM-based evaluator. In contrast, PhyI2V’s high score in (d) stems from a valid advantage in world knowledge (litmus solution chemistry).