cs.CVOct 1, 2026

HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

Authors: Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag

Organizations: Virginia Tech · Qualcomm AI Research

Abstract

Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. NEWTON: Agentic Planning for Physically Grounded Video Generation

    May 18, 2026Yuxiang Feng, Juncheng Wang, Chao Xu +7Video GenerationIterative Re-Planning

  2. PhyWorld: Physics-Faithful World Model for Video Generation

    May 19, 2026Pu Zhao, Juyi Lin, Timothy Rupprecht +10Generative Video ModelsVideo Generation

  3. PhyCo: Learning Controllable Physical Priors for Generative Motion

    Apr 30, 2026Sriram Narayanan, Ziyu Jiang, Srinivasa Narasimhan +1Generative Video ModelsVideo Diffusion Models