Organizations: School of Computer Science, Nanjing University. · SenseTime. · Shenzhen Technology University. · Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology.
Generating physically plausible videos for solid-gas dynamics is challenging as different phases exhibit distinct dynamics yet remain coupled through physical interactions. We present PAVG, a Phase-Aware Video Generator for solid-gas dynamics and interactions. It employs a dual-branch architecture to explicitly model the distinct dynamics of solids and gases, while spatiotemporal cross-attention captures their physical interactions. This design enables PAVG to preserve phasespecific motion characteristics while producing physically consistent responses across phases. To facilitate this task, we further construct a simulation corpus comprising over 700K physical trajectories across diverse solid, gas, and solid-gas interaction scenarios. Extensive evaluations demonstrate that our PAVG produces videos with improved motion adherence, physical plausibility, and visual quality compared with existing approaches.
Figures & tables
Figure 1: A tennis ball moving through dust. Wan2.2 moves the ball in the wrong direction; Force Prompting leaves the dust stationary; FlashMotion moves the ball and dust together without a visible interaction. PAVG follows the prescribed ball motion and produces a responsive dust plume, yielding a more physically plausible solid–gas interaction. The figure contains fine-grained motion details and is best viewed on a computer.
Figure 2: PAVG overview: (a) initialize phase-specific 3D point clouds; (b) use a phase-aware point-trajectory predictor to predict solid and gas trajectories; (c) project the trajectories to guide a pretrained video generation model. Flames and snowflakes mark trainable and frozen modules.
Figure 3: PAVG architecture: (a) phase-specific denoising with cross-phase interaction (CPI) modules; (b) gated cross-attention within CPI; and (c) spatial and temporal attention axes. Hg′ denotes the gas latent feature after the CPI update.
Figure 4
Method
Overall ∗
Single-phase: Solid
Single-phase: Gas
Interaction: Solid–gas
vIoU ↑
CD ↓
Corr ↓
vIoU ↑
CD ↓
Corr ↓
vIoU ↑
CD ↓
Corr ↓
vIoU ↑
CD ↓
Corr ↓
Gas-only
SFBC
—
—
0.483874
0.120906
0.274170
—
DLF
—
—
0.483505
0.053650
0.241004
—
SEGNN
—
—
0.561893
0.007060
0.176081
—
Solid-only
Table 2: Trajectory prediction across single-phase motion and solid–gas interaction.
Solid
Predictor
vIoU ↑
CD ↓
Corr ↓
Solid-only
0.588372
0.022882
0.080403
Gas-only
N/A
N/A
N/A
Unified
0.571150
0.024851
0.086428
Table 3: Phase-specific versus unified modeling in Stage 1. Solid-only and Gas-only are trained separately; Unified shares one predictor across both phases. Training exposure is matched per phase. Solid scores aggregate the five solid subsets.
Figure 5: Qualitative interaction-pathway ablation on three held-out solid–gas cases. With the CPI pathway enabled, the gas visibly responds to and follows the solid’s motion.
Variant
Joint solid–gas
Interaction: solid
Interaction: gas
vIoU ↑
CD ↓
Corr ↓
vIoU ↑
CD ↓
Corr ↓
vIoU ↑
CD ↓
Corr ↓
PAVG
0.690028
0.003928
0.059928
0.728056
0.011593
0.049734
0.642296
0.003537
0.070121
w/o solid-to-gas Cross-Attention
0.540755
0.016072
0.178830
0.728056
0.011593
0.049734
0.428435
0.067136
0.307926
Table 4: Quantitative interaction-pathway ablation on the solid–gas interaction subset of our benchmark. The same PAVG checkpoint is evaluated with or without solid-to-gas Cross-Attention. Metrics are reported for the concatenated joint output and separately for each phase.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Simulator
Train
Test
Solid
elastic-force
MPM
150,201
100
elastic-gravity
MPM
94,038
100
plasticine-gravity
MPM
71,356
100
sand-gravity
MPM
95,414
100
rigid-gravity
rigid body
97,091
100
Appendix
Table 5: Training datasets and held-out trajectory test sets. The force condition denotes drag for solid sequences and initial velocity for gas sequences.
Setting
Value
Blocks / latent dimension / attention heads
8 / 256 / 4 per branch
Points / future frames
2,048 per phase / 24
Optimizer
AdamW, β=(0.9,0.999) , ϵ=10−8
Learning rate / weight decay
10−4 / 10−2
Schedule / warmup
Cosine decay / 100 updates
Gradient clipping / precision
Norm 1.0 / bfloat16
Appendix
Table 6: PAVG architecture and optimization settings.
Initialization
Overall ∗
Single-phase: Solid
Single-phase: Gas
Interaction: Solid–gas
vIoU ↑
CD ↓
Corr ↓
vIoU ↑
CD ↓
Corr ↓
vIoU ↑
CD ↓
Corr ↓
vIoU ↑
CD ↓
Corr ↓
Random
0.644972
0.008163
0.082704
0.582268
0.022199
0.080927
0.710286
0.002036
0.096300
0.643667
0.004209
0.076795
Phase-specific
0.690964
0.006586
0.063135
0.617950
0.016784
0.067902
0.783433
0.001444
0.062611
0.681236
0.004059
0.061014
Appendix
Table 7: Stage 2 initialization ablation under the same dual-branch architecture and reduced-loss training recipe. Overall uses the weighting in Table 2 .
Objective
vIoU ↑
CD ↓
Corr ↓
w/o Lcoupling
0.677608
0.003948
0.061509
with Lcoupling
0.681236
0.004059
0.061014
Appendix
Table 8: Ablation of the cross-phase interaction physical constraint on the solid–gas interaction subset of our benchmark.
GPT-4o
GPT-5.5
GPT-5.6-sol
Gemini- 3.8-flash
DeepSeek- v4.1-flash
Method
SA ↑
PC ↑
VQ ↑
SA ↑
PC ↑
VQ ↑
SA ↑
PC ↑
VQ ↑
SA ↑
PC ↑
VQ ↑
SA ↑
PC ↑
VQ ↑
HunyuanVideo-1.5
3.63
3.63
3.88
3.06
2.91
3.59
2.94
2.75
3.69
3.00
3.00
3.63
2.94
2.88
3.53
Wan2.2-I2V-A14B
3.31
3.50
3.81
3.16
3.19
3.44
2.53
2.41
3.34
2.94
3.00
3.53
3.31
3.28
3.53
CogVideoX1.5-5B-I2V
2.75
2.94
3.41
2.13
2.31
2.72
1.84
2.19
2.75
2.31
2.47
2.91
2.25
2.41
2.84
DragAnything
2.78
2.81
3.06
2.38
2.31
2.50
1.94
2.09
2.41
2.34
2.44
2.72
2.56
2.59
2.78
ObjCtrl-2.5D
2.59
2.66
3.03
2.09
2.03
2.75
2.16
2.06
2.75
2.50
2.53
2.88
2.03
2.06
2.72
Appendix
Table 9: Individual judge scores over the same 32 scenes. Higher is better for all metrics. Bold : best; underlined : second best within each judge and metric.
We present Fluid-Gen-Zero, a training-free framework for physics-aware fluid-object interaction video generation that decouples physical reasoning from appearance synthesis. Our key insight is to delegate motion dynamics to a physics simulator while preserving the appearance modeling capacity of pretrained video generators. We bridge these two domains through a two-level agentic workflow: generation-time planning, where a vision-language model (VLM) agent interprets intent and the simulation rollout to organize generation clips, and latent-space guidance, which injects simulation signals into denoising through region-aware latent wrapping. This plug-and-play design is compatible with current video foundation models. We further introduce a benchmark for fluid-object interaction video generation. Across Tora (CogVideoX-based), VACE and WanMove (Wan-based), Fluid-Gen-Zero consistently improves simulation alignment, reducing object trajectory error by 26.7%-81.5% and fluid fEPE (fluid flow endpoint error) by 67.9%-84.0%, while largely preserving perceptual quality. In a human preference study, raters favor Fluid-Gen-Zero in 55.1%-74.4% of same-backbone comparisons across three backbones, and in 90.4%-94.2% of comparisons against simulation-based methods. Code and data will be released upon acceptance.
Hong Huang, Yuqiu Liu, Chenyu You +3
Simon Fraser University · Stony Brook University · Lawrence Berkeley National Laboratory +1
Recent advances in physics-grounded video generation leverage physics simulation as a physical prior to guide video synthesis toward physically plausible outcomes. The simulation process is controlled by physical specifications, which are typically generated by a vision-language model in a single pass. Such one-shot prediction often fails to accurately translate user intent into executable simulations, particularly for fine-grained object dynamics, complex motion trajectories, and temporally structured interactions. In this paper, we propose PhysAgent, a reflective agentic framework that closes the loop among physical program generation, physics simulation, stage-specific verification, and targeted program repair. Beyond improving the control of coupled physical parameters, our framework enables the agent to progressively realize complex trajectories, multi-stage interactions, and precise event outcomes by treating each physical program as an executable hypothesis. In addition, we design a set of physics-control APIs to support more stable and complex motion behaviors. Extensive experiments demonstrate that PhysAgent produces more physically plausible videos, achieves better prompt alignment, and generalizes more effectively across diverse physical scenarios.
Qirui Li, Jinkun Hao, Yibo Li +3
Shanghai Jiao Tong University · Cardiff University
While recent video generation models have achieved significant visual fidelity, they often suffer from the lack of explicit physical controllability and plausibility. To address this, some recent studies attempted to guide the video generation with physics-based rendering. However, these methods face inherent challenges in accurately modeling complex physical properties and effectively control ling the resulting physical behavior over extended temporal sequences. In this work, we introduce PhysChoreo, a novel framework that can generate videos with diverse controllability and physical realism from a single image. Our method consists of two stages: first, it estimates the static initial physical properties of all objects in the image through part-aware physical property reconstruction. Then, through temporally instructed and physically editable simulation, it synthesizes high-quality videos with rich dynamic behaviors and physical realism. Experimental results show that PhysChoreo can generate videos with rich behaviors and physical realism, outperforming state-of-the-art methods on multiple evaluation metrics.