Fluid-Gen-Zero: Grounding Pretrained Video Generators in Physics without Training
Organizations: Simon Fraser University · Stony Brook University · Lawrence Berkeley National Laboratory · Meta Reality Labs
Abstract
We present Fluid-Gen-Zero, a training-free framework for physics-aware fluid-object interaction video generation that decouples physical reasoning from appearance synthesis. Our key insight is to delegate motion dynamics to a physics simulator while preserving the appearance modeling capacity of pretrained video generators. We bridge these two domains through a two-level agentic workflow: generation-time planning, where a vision-language model (VLM) agent interprets intent and the simulation rollout to organize generation clips, and latent-space guidance, which injects simulation signals into denoising through region-aware latent wrapping. This plug-and-play design is compatible with current video foundation models. We further introduce a benchmark for fluid-object interaction video generation. Across Tora (CogVideoX-based), VACE and WanMove (Wan-based), Fluid-Gen-Zero consistently improves simulation alignment, reducing object trajectory error by 26.7%-81.5% and fluid fEPE (fluid flow endpoint error) by 67.9%-84.0%, while largely preserving perceptual quality. In a human preference study, raters favor Fluid-Gen-Zero in 55.1%-74.4% of same-backbone comparisons across three backbones, and in 90.4%-94.2% of comparisons against simulation-based methods. Code and data will be released upon acceptance.
Figures & tables
| Method | Size | Obj. Traj. Err. | Fluid fEPE | Subj. Cons. | Temp. Flick. | Aesth. Qual. | Img. Qual. | Physical Commonsense | GPT Judge | |||
| PhysR ( ) | PhotoR ( ) | Sem ( ) | Overall ( ) | |||||||||
| PhysGen3D | - | - | - | 0.873 | 0.998 | 0.447 | 0.654 | 4.791 | 0.260 0.10 | 0.517 0.13 | 0.453 0.21 | 0.410 |
| WonderPlay | 5B | - | - | 0.731 | 0.991 | 0.362 | 0.435 | 4.558 | 0.298 0.09 | 0.421 0.09 | 0.398 0.15 | 0.372 |
| MotionCraft | 1B | 169.62 | 10.659 | 0.836 | 0.977 | 0.438 | 0.598 | 4.814 | 0.315 0.15 | 0.475 0.19 | 0.475 0.21 | 0.421 |
| VACE | 1.3B | 56.02 | 6.445 | 0.733 | 0.978 | 0.461 | 0.590 | 4.116 | 0.385 0.12 | 0.528 0.14 | 0.482 0.19 | 0.465 |
| + Fluid-Gen-Zero | 1.3B | 41.06 | 2.070 | 0.782 | 0.990 | 0.486 | 0.642 | 4.326 | 0.425 0.11 | 0.613 0.10 | 0.578 0.21 | 0.539 |
| Comparison | Sim. Align. | Temp. Cons. | Overall Pref. |
| Same-backbone comparison | |||
| Ours (WanMove) vs. WanMove | 55.1% | 60.3% | 55.1% |
| Ours (Tora) vs. Tora | 69.2% | 69.2% | 74.4% |
| Ours (VACE) vs. VACE | 66.7% | 69.2% | 69.2% |
| Simulation-based methods | |||
| Fluid-Gen-Zero vs. PhysGen3D | 95.2% | 85.6% | 94.2% |
| Symbol | Signal | Definition at transition |
| Physical family : 3D particle trajectories, computed per state | ||
| Centroid speed | . | |
| Mean particle speed | Mean of over particles . | |
| Particle-speed variance | Variance of the same displacement magnitudes over . | |
| High-percentile displacement | th percentile of the displacement magnitudes over . | |
| Particle acceleration | Mean magnitude of the finite difference of displacement vectors. | |
| Event type | Alignment set | Signal names |
| impact / contact | contact intensity, high-percentile flow, silhouette change | |
| splash | high-percentile fluid displacement, flow discontinuity, fragmentation | |
| split / merge | fragmentation, fluid spread change, flow discontinuity | |
| swirl | flow curl, flow-magnitude variance, fluid particle-speed variance | |
| occlusion | boundary density, silhouette change, contact change | |
| material detail | mean flow magnitude, boundary density |
| Benchmark / Data | Statistics | Annotations | ||||
| Samples | Frames | Categorization | Mask | 3D Points | Optical Flow | |
| MagicBench [ 21 ] | 600 | 49 | ||||
| MoveBench [ 11 ] | 1,018 | 81 | ||||
| RealWonder [ 26 ] | 30 | – | ||||
| PerpetualWonder [ 49 ] | 10 | – | ||||
| Ours | 43 | 401 (simulation rollout) | ||||
| Method | Size | Obj. Traj. Err. | Fluid fEPE | Subj. Cons. | Temp. Flick. | Aesth. Qual. | Img. Qual. | Physical Commonsense |
| Tora | 5B | 133.20 | 3.950 | 0.804 | 0.980 | 0.487 | 0.656 | 4.512 |
| Tora w/ Planning | 5B | 128.12 | 3.221 | 0.816 | 0.982 | 0.487 | 0.653 | 4.605 |
| Method | Size | Obj. Traj. Err. | Fluid fEPE | Subj. Cons. | Temp. Flick. | Aesth. Qual. | Img. Qual. | Physical Commonsense |
| Tora | 5B | 133.20 | 3.950 | 0.804 | 0.980 | 0.487 | 0.656 | 4.512 |
| Tora w/ Latent-Space Guidance | 5B | 24.79 | 0.812 | 0.912 | 0.995 | 0.478 | 0.653 | 4.628 |
| Setting | Obj. Traj. Err. | Fluid fEPE | Subj. Cons. | Temp. Flick. | Aesth. Qual. | Img. Qual. | Physical Commonsense |
| Full | 24.61 | 0.843 | 0.906 | 0.995 | 0.506 | 0.656 | 4.628 |
| w/o Region-Aware Wrapping | 30.40 | 1.026 | 0.882 | 0.985 | 0.476 | 0.657 | 4.628 |
| w/o Fluid Wrapping | 28.06 | 0.937 | 0.897 | 0.998 | 0.474 | 0.652 | 4.605 |