Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.
Figures & tables
Figure 1: HiPhy enables generation of physically plausible videos for scenes with multiple concurrent physical principles. Wan (left) collapses on multi-principle prompts, failing to generate coherent motion or dropping a physical principle entirely, such as missing boiling or steam motions. Instead of collapsing physical plausibility into a single score, HiPhy (right) aligns generation with a hierarchical reward that tracks the temporal sub-stages of each principle as an independent learning stream, jointly recovering all concurrent processes with correct temporal progression.
Figure 2: HiPhy explicitly decomposes complex physical events (e.g., an apple falling into water) into temporally ordered sub-stages. This allows us to recursively score the chronological progression, completeness, and alignment of the underlying physics.
Figure 3: Overview of the HiPhy framework. We introduce a hierarchical optimization framework aimed to elicit physically grounded video generation in scenes with multiple physical principles.
Figure 4: Qualitative results. HiPhy (top row) produces coherent motion and correct temporal progression on prompts spanning single- and multi-principle physical scenes, while Wan2.1 (bottom row) frequently drops a principle, breaks contact dynamics, or generates static scenes.
Figure 5: Qualitative comparisons. As the number of physical principles per prompt grows, baselines exhibit consistent failure modes, incoherent motion, hallucinated objects, or skipping a principle entirely, and these compound on three-principle prompts (bottom) where multiple processes are dropped simultaneously. HiPhy (top row of each block) renders each constituent process coherently without degrading visual quality.
VideoPhy2
PhyWorldBench
MultiPhyBench
VBench
Method
PC ↑
SA ↑
PC & SA ↑
PC ↑
SA ↑
PC & SA ↑
PC ↑
SA ↑
PC & SA ↑
Dyn. ↑
App. ↑
Wan2.1
49.10
43.60
30.5
47.7
40.9
23.7
53.0
35.0
27.0
0.808
0.573
PnP
80.1
68.0
60.14
56.1
48.5
31.2
57.0
37.0
29.0
0.844
0.549
VideoRepa
70.9
53.5
43.8
51.7
33.3
19.3
49.0
22.0
15.0
0.863
0.550
PhyT2V
79.0
65.0
55.0
40.3
47.5
22.7
50.0
26.0
15.0
0.872
0.500
WISA
68.0
59.7
44.3
56.3
50.7
30.7
52.5
43.0
24.5
0.859
0.567
Table 1: Comprehensive evaluation across VideoPhy2, PhyWorldBench, Multi-concept, and VBench, and ablation experiments over different reward terms
Figure 7: Performance scaling and human evaluation. (A–C) Automatic metrics for PC, SA, and joint PC&SA across process counts. HiPhy (purple) maintains a significantly wider margin as complexity increases. (D) Averaged human scores per principle (P1–P4). While baselines show constant degradation in multi-event scenes, HiPhy achieves consistent physical fidelity across all concurrent principles.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Physics Domain
#Videos
Stage Quality
Ordering
Overall
mean ± 95% CI
mean ± 95% CI
mean
Rigid-body / collision
11
4.42 ± 0.21
4.61 ± 0.23
4.46
Fluid dynamics
12
4.31 ± 0.23
4.55 ± 0.25
4.36
Thermodynamics
8
4.05 ± 0.30
4.28 ± 0.32
4.10
Deformation
5
4.24 ± 0.34
4.40 ± 0.36
4.27
Optics
8
3.88 ± 0.32
4.12 ± 0.34
3.92
Appendix
Table 4: Human validation of extracted sub-stages across n=50 videos and 5 raters. Each video was rated by all 5 raters on a 1–5 Likert scale. We report mean ratings with 95% confidence intervals for individual sub-stage rendering quality (averaged across all sub-stages of a video) and temporal ordering.
Physical correctness
Ordering
Completeness
Extracted tree
4.12
3.78
3.86
Corrupted tree
2.23
1.80
2.40
Appendix
Table 5: Human evaluation results for extracted versus corrupted stage trees.
Benchmark
Variant
PC
SA
VideoPhy2
Hierarchical
66.0
52.0
VideoPhy2
Flat
63.0
47.0
MultiPhyBench
Hierarchical
59.0
44.0
MultiPhyBench
Flat
47.0
29.0
Appendix
Table 6: Ablation on reward structure
Reward Variant
Omission rate (%) ↓
Ordering accuracy (%) ↑
Stage match (%) ↑
Full
6.1
86.0
81.0
( − ) ordering
6.8
61.0
72.0
( − ) completeness
17.4
84.0
79.0
( − ) alignment
9.7
83.0
68.0
Appendix
Table 7: Gemini–human agreement is computed as the fraction of items on which the automatic and human labels match: 0.84 (omission), 0.77 (ordering), 0.82 (stage match).
Exp3
PC (N=1)
PC (N=2)
PC (N=3)
Overall SA
Full
72.1
63.0
48.0
58.3
w/o per-stream norm
70.3
38.0
22.3
21.6
w/o 1/n rescaling
72.8
41.3
37.6
35.0
Appendix
Table 8: Ablation results on normalization and rescaling components.
Figure 8: User study sample question
VideoPhy2
PhyWorldBench
Multi-concept
VBench
Method
PC ↑
SA ↑
Both ↑
PC ↑
SA ↑
Both ↑
PC ↑
SA ↑
Both ↑
Dyn. ↑
App. ↑
Wan2.1
49.10 ± 2.06
43.60 ± 2.04
30.50 ± 1.90
47.7 ± 2.91
40.9 ± 2.86
23.7 ± 2.48
53.0 ± 2.88
35.0 ± 2.75
27.0 ± 2.56
0.808 ± 0.014
0.573 ± 0.003
PnP
80.1 ± 1.64
68.0 ± 1.92
60.14 ± 2.02
56.1 ± 2.89
48.5 ± 2.91
31.2 ± 2.70
57.0 ± 2.86
37.0 ± 2.79
29.0 ± 2.62
0.844 ± 0.013
0.549 ± 0.003
VideoRepa
70.9 ± 1.87
53.5 ± 2.05
43.8 ± 2.04
51.7 ± 2.91
33.3 ± 2.74
19.3 ± 2.30
49.0 ± 2.89
22.0 ± 2.39
15.0 ± 2.06
0.863 ± 0.012
0.550 ± 0.003
PhyT2V
79.0 ± 1.68
65.0 ± 1.96
55.0 ± 2.05
40.3 ± 2.86
47.5 ± 2.91
22.7 ± 2.44
50.0 ± 2.89
26.0 ± 2.53
15.0 ± 2.06
0.872 ± 0.012
0.500 ± 0.003
WISA
68.0 ± 1.92
59.7 ± 2.02
44.3 ± 2.05
56.3 ± 2.89
50.7 ± 2.91
30.7 ± 2.69
52.5 ± 2.88
43.0 ± 2.86
24.5 ± 2.48
0.859 ± 0.012
0.567 ± 0.003
Appendix
Table 9: Comprehensive evaluation across VideoPhy2, PhyWorldBench, Multi-concept, and VBench along with their standard deviations.