Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.
Figures & tables
Figure 1: HiPhy enables generation of physically plausible videos for scenes with multiple concurrent physical principles. Wan (left) collapses on multi-principle prompts, failing to generate coherent motion or dropping a physical principle entirely, such as missing boiling or steam motions. Instead of collapsing physical plausibility into a single score, HiPhy (right) aligns generation with a hierarchical reward that tracks the temporal sub-stages of each principle as an independent learning stream, jointly recovering all concurrent processes with correct temporal progression.
Figure 2: HiPhy explicitly decomposes complex physical events (e.g., an apple falling into water) into temporally ordered sub-stages. This allows us to recursively score the chronological progression, completeness, and alignment of the underlying physics.
Figure 3: Overview of the HiPhy framework. We introduce a hierarchical optimization framework aimed to elicit physically grounded video generation in scenes with multiple physical principles.
Figure 4: Qualitative results. HiPhy (top row) produces coherent motion and correct temporal progression on prompts spanning single- and multi-principle physical scenes, while Wan2.1 (bottom row) frequently drops a principle, breaks contact dynamics, or generates static scenes.
Figure 5: Qualitative comparisons. As the number of physical principles per prompt grows, baselines exhibit consistent failure modes, incoherent motion, hallucinated objects, or skipping a principle entirely, and these compound on three-principle prompts (bottom) where multiple processes are dropped simultaneously. HiPhy (top row of each block) renders each constituent process coherently without degrading visual quality.
VideoPhy2
PhyWorldBench
MultiPhyBench
VBench
Method
PC ↑
SA ↑
PC & SA ↑
PC ↑
SA ↑
PC & SA ↑
PC ↑
SA ↑
PC & SA ↑
Dyn. ↑
App. ↑
Wan2.1
49.10
43.60
30.5
47.7
40.9
23.7
53.0
35.0
27.0
0.808
0.573
PnP
80.1
68.0
60.14
56.1
48.5
31.2
57.0
37.0
29.0
0.844
0.549
VideoRepa
70.9
53.5
43.8
51.7
33.3
19.3
49.0
22.0
15.0
0.863
0.550
PhyT2V
79.0
65.0
55.0
40.3
47.5
22.7
50.0
26.0
15.0
0.872
0.500
WISA
68.0
59.7
44.3
56.3
50.7
30.7
52.5
43.0
24.5
0.859
0.567
Table 1: Comprehensive evaluation across VideoPhy2, PhyWorldBench, Multi-concept, and VBench, and ablation experiments over different reward terms
Figure 7: Performance scaling and human evaluation. (A–C) Automatic metrics for PC, SA, and joint PC&SA across process counts. HiPhy (purple) maintains a significantly wider margin as complexity increases. (D) Averaged human scores per principle (P1–P4). While baselines show constant degradation in multi-event scenes, HiPhy achieves consistent physical fidelity across all concurrent principles.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Physics Domain
#Videos
Stage Quality
Ordering
Overall
mean ± 95% CI
mean ± 95% CI
mean
Rigid-body / collision
11
4.42 ± 0.21
4.61 ± 0.23
4.46
Fluid dynamics
12
4.31 ± 0.23
4.55 ± 0.25
4.36
Thermodynamics
8
4.05 ± 0.30
4.28 ± 0.32
4.10
Deformation
5
4.24 ± 0.34
4.40 ± 0.36
4.27
Optics
8
3.88 ± 0.32
4.12 ± 0.34
3.92
Appendix
Table 4: Human validation of extracted sub-stages across n=50 videos and 5 raters. Each video was rated by all 5 raters on a 1–5 Likert scale. We report mean ratings with 95% confidence intervals for individual sub-stage rendering quality (averaged across all sub-stages of a video) and temporal ordering.
Physical correctness
Ordering
Completeness
Extracted tree
4.12
3.78
3.86
Corrupted tree
2.23
1.80
2.40
Appendix
Table 5: Human evaluation results for extracted versus corrupted stage trees.
Benchmark
Variant
PC
SA
VideoPhy2
Hierarchical
66.0
52.0
VideoPhy2
Flat
63.0
47.0
MultiPhyBench
Hierarchical
59.0
44.0
MultiPhyBench
Flat
47.0
29.0
Appendix
Table 6: Ablation on reward structure
Reward Variant
Omission rate (%) ↓
Ordering accuracy (%) ↑
Stage match (%) ↑
Full
6.1
86.0
81.0
( − ) ordering
6.8
61.0
72.0
( − ) completeness
17.4
84.0
79.0
( − ) alignment
9.7
83.0
68.0
Appendix
Table 7: Gemini–human agreement is computed as the fraction of items on which the automatic and human labels match: 0.84 (omission), 0.77 (ordering), 0.82 (stage match).
Exp3
PC (N=1)
PC (N=2)
PC (N=3)
Overall SA
Full
72.1
63.0
48.0
58.3
w/o per-stream norm
70.3
38.0
22.3
21.6
w/o 1/n rescaling
72.8
41.3
37.6
35.0
Appendix
Table 8: Ablation results on normalization and rescaling components.
Figure 8: User study sample question
VideoPhy2
PhyWorldBench
Multi-concept
VBench
Method
PC ↑
SA ↑
Both ↑
PC ↑
SA ↑
Both ↑
PC ↑
SA ↑
Both ↑
Dyn. ↑
App. ↑
Wan2.1
49.10 ± 2.06
43.60 ± 2.04
30.50 ± 1.90
47.7 ± 2.91
40.9 ± 2.86
23.7 ± 2.48
53.0 ± 2.88
35.0 ± 2.75
27.0 ± 2.56
0.808 ± 0.014
0.573 ± 0.003
PnP
80.1 ± 1.64
68.0 ± 1.92
60.14 ± 2.02
56.1 ± 2.89
48.5 ± 2.91
31.2 ± 2.70
57.0 ± 2.86
37.0 ± 2.79
29.0 ± 2.62
0.844 ± 0.013
0.549 ± 0.003
VideoRepa
70.9 ± 1.87
53.5 ± 2.05
43.8 ± 2.04
51.7 ± 2.91
33.3 ± 2.74
19.3 ± 2.30
49.0 ± 2.89
22.0 ± 2.39
15.0 ± 2.06
0.863 ± 0.012
0.550 ± 0.003
PhyT2V
79.0 ± 1.68
65.0 ± 1.96
55.0 ± 2.05
40.3 ± 2.86
47.5 ± 2.91
22.7 ± 2.44
50.0 ± 2.89
26.0 ± 2.53
15.0 ± 2.06
0.872 ± 0.012
0.500 ± 0.003
WISA
68.0 ± 1.92
59.7 ± 2.02
44.3 ± 2.05
56.3 ± 2.89
50.7 ± 2.91
30.7 ± 2.69
52.5 ± 2.88
43.0 ± 2.86
24.5 ± 2.48
0.859 ± 0.012
0.567 ± 0.003
Appendix
Table 9: Comprehensive evaluation across VideoPhy2, PhyWorldBench, Multi-concept, and VBench along with their standard deviations.
Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We identify a specification bottleneck: text prompts are lossy compression of the physical world, omitting the parameters that fully determine dynamics, and no amount of model scaling can recover what was never specified. From this diagnosis we derive three properties that physics conditioning must satisfy -- sufficiency, dynamism, and verifiability -- and show that no existing approach satisfies all three. We present NEWTON, in which video generation is demoted from the system output to one action inside an agent's toolbox: a learned planner orchestrates physics-aware tools (keyframe generation, scientific computation, prompt refinement) to construct rich conditioning, and a verifier closes the loop for iterative re-planning. The planner is the sole trainable component, optimized on-policy via Flow-GRPO inside the live multi-turn loop. On VideoPhy-2, NEWTON improves joint accuracy from 21.4% to 29.7% on LTX-Video and from 30.7% to 37.4% on Veo-3.1, without modifying either generator. Our project page: https://Newton026.github.io/newton
Yuxiang Feng, Juncheng Wang, Chao Xu +7
Zhejiang University · The Hong Kong Polytechnic University · IROOTECH TECHNOLOGY +1
World simulators can provide safe and scalable environments for training Physical AI systems before real-world deployment. Large video generation models are emerging as a promising basis for such simulators because they can generate diverse and realistic visual futures. However, using them as world simulators requires physically faithful video continuations, namely, generated videos that preserve the physical state implied by the conditioning input, and evolve in ways consistent with basic physical principles. We propose PhyWorld, a video generation world model designed to produce temporally coherent and physically faithful scene continuations through two-stage post-training. In the first stage, we improve video-to-video continuation with flow matching fine-tuning, encouraging stable visual attributes and coherent motion dynamics across frames. In the second stage, we align generated dynamics with physical principles using Direct Preference Optimization (DPO) over physics preference pairs, guiding the model toward outputs with higher physical plausibility. To evaluate PhyWorld, we use both standard video-quality benchmarks and a dedicated physical-faithfulness benchmark with per-law scoring. Experiments show that PhyWorld improves video consistency, achieving an average score of 0.769 on VBench compared with 0.756 or below for state-of-the-art baselines. PhyWorld also improves physical plausibility, reaching an average score of 3.09 on our physical-faithfulness benchmark compared with 2.99 for the strongest baseline. These results suggest that post-training large video generation models with continuation and physics-preference signals can make them more effective world simulators for Physical AI.
Pu Zhao, Juyi Lin, Timothy Rupprecht +10
Northeastern University · University of Georgia · Tulane University +1
Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties. We present PhyCo, a framework that introduces continuous, interpretable, and physically grounded control into video generation. Our approach integrates three key components: (i) a large-scale dataset of over 100K photorealistic simulation videos where friction, restitution, deformation, and force are systematically varied across diverse scenarios; (ii) physics-supervised fine-tuning of a pretrained diffusion model using a ControlNet conditioned on pixel-aligned physical property maps; and (iii) VLM-guided reward optimization, where a fine-tuned vision-language model evaluates generated videos with targeted physics queries and provides differentiable feedback. This combination enables a generative model to produce physically consistent and controllable outputs through variations in physical attributes-without any simulator or geometry reconstruction at inference. On the Physics-IQ benchmark, PhyCo significantly improves physical realism over strong baselines, and human studies confirm clearer and more faithful control over physical attributes. Our results demonstrate a scalable path toward physically consistent, controllable generative video models that generalize beyond synthetic training environments.