PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos
Organizations: NVIDIA · University of Waterloo · Vector Institute
Abstract
Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can often be myopic to physical dynamics, or fine-tuned evaluators trained on human annotations, which overfit to dataset-specific cues and fail to generalize. A key challenge is that existing supervision sources provide either relative ordering or absolute scores, but not both reliably and consistently across varied settings. To this end, we introduce PhyProbe, an evaluator that extracts features from a frozen pretrained spatio-temporal encoder and maps them to a scalar physical consistency violation score via a lightweight scoring head. PhyProbe is trained through a unified objective combining pairwise ranking, regression on noisy scalar annotations, and anchor-based calibration over a curated set of heterogeneous supervision sources. Experiments show that PhyProbe outperforms prior methods on most pairwise benchmarks spanning real-generated and generated-generated pairs under varying correspondence, with the largest gains in no-correspondence and generated-generated settings where existing fine-tuned evaluators degrade sharply. PhyProbe achieves strong correlation with human judgments, with close agreement between rank-based and linear metrics, indicating that scores are both well ordered and anchored to a stable [0, 1] scale. Further, despite being trained on supervision indicative of physical consistency, without explicit general-preference labels, PhyProbe also performs competitively on human preference benchmarks: consistent with the observation that physics violations are entangled with broader quality degradations.
Figures & tables
| Training Dataset | Pairs | Type | Correspondence | Rank | Mag | Cal |
|---|---|---|---|---|---|---|
| PAI-Bench [ 60 ] | 42K | R-G | Image + Text | ✓ | ✓ | |
| VideoPhy2 [ 6 ] | 19K | G-G | Action / Text | ✓ | ✓ | |
| VideoFeedback2 [ 19 ] | 31K | G-G | None | ✓ | ✓ | |
| BrokenVideos [ 28 ] | 14K | G-G | None | ✓ | ✓ | |
| ImpossibleVideos [ 2 ] | 27K | R-G | None | ✓ | ✓ | |
| TRAVL [ 36 ] | 98K | R-G | None | ✓ | ✓ |
| Pairwise Accuracy (%) | Human Correlation | ||||||||||||
| Img+Txt | Text | None | Avg. | VP2 | VF2 | Avg. | |||||||
| Method | Backbone | ImplB p | VP2 p | PhyDEx p | VF2 p | ||||||||
| (R–G) | (G–G) | (R–G) | (G–G) | ||||||||||
| VP2-AutoEval | VideoCon (7B) | 28.0 | 28.5 | 30.5 | 36.3 | 32.1 | 0.359 | 0.362 | 0.235 | 0.240 | 0.298 | 0.302 | |
| V-JEPA Surprise | V-JEPA2-H/16 (600M) | 33.3 | 46.3 | 49.0 | 45.5 | 46.6 | 0.065 | 0.065 | 0.121 | 0.133 | 0.093 | 0.099 | |
| VideoScore2 | Qwen2.5-VL (7B) | 68.7 | 49.7 | 60.6 | 65.0 | 59.9 | 0.197 | 0.160 | 0.462 | 0.430 | 0.336 | 0.301 | |
| Method | MonetBench | Rapidata-I2V | VideoGen-RewardBench | ||||||
|---|---|---|---|---|---|---|---|---|---|
| w/ tie | w/o tie | w/o tie only | Overall | MQ | VQ | ||||
| w/ tie | w/o tie | w/ tie | w/o tie | w/ tie | w/o tie | ||||
| Random | 33.2 | 49.9 | 49.7 | 33.2 | 49.7 | 33.2 | 49.5 | 33.3 | 49.8 |
| V-JEPA Surprise | 52.4 | 53.8 | 45.1 | 45.2 | 44.2 | 48.8 | 47.1 | 48.7 | 47.6 |
| VideoScore2 | 42.3 | 41.1 | 45.8 | 57.4 | 58.9 | 54.3 | 60.6 | 55.3 | 60.1 |
| VideoPhy-2-AutoEval | 22.5 | 16.5 | 9.5 | 24.6 | 19.5 | 37.9 | 20.5 | 34.4 | 20.4 |
| Variant | Pairwise Accuracy (%) | Human Correlation | Score | ||||||
| None | Txt | Img+Txt | VP2 | VF2 | Range | ||||
| PhyDEx p | VF2 p | VP2 p | ImplB p | ||||||
| (R–G) | (G–G) | (G–G) | (R–G) | ||||||
| PhyProbe ( ) | 94.7 | 75.6 | 61.9 | 99.2 | 0.311 | 0.316 | 0.614 | 0.609 | |
| PhyProbe ( ) | 98.7 | 75.3 | 62.3 | 94.1 | 0.315 | 0.316 | 0.584 | 0.583 | |
| PhyProbe ( ) | 68.7 | 79.3 | 64.4 | 84.9 | 0.363 | 0.364 | 0.664 | 0.656 | |
| Pairwise Accuracy (%) | Human Correlation | |||||||||
| Img+Txt | Text | None | VP2 | VF2 | ||||||
| PhyProbe | Params | ImplB p | VP2 p | PhyDEx p | VF2 p | |||||
| Backbone | (R–G) | (G–G) | (R–G) | (G–G) | ||||||
| Wan2.2 VAE | 300M | 71.7 | 58.3 | 67.1 | 74.3 | 0.186 | 0.194 | 0.562 | 0.581 | |
| V-JEPA2-VIT-L/16 | 300M | 76.1 | 60.6 | 88.8 | 74.2 | 0.228 | 0.206 | 0.564 | 0.580 | |
| V-JEPA2-VIT-H/16 | 600M | 81.4 | 60.2 | 70.7 | 76.3 | 0.301 | 0.303 | 0.623 | 0.630 | |
| Head | Trainable | Input | Rel. (×) | VideoPhy2 | VideoFeedback2 | ||
|---|---|---|---|---|---|---|---|
| Params | Dim | Acc | Acc | ||||
| Mean (MLP) | 0.36M | 1 | 61.0 | 0.277 | 73.2 | 0.578 | |
| Attention Pooling | 6.92M | 19.2 | 63.7 | 0.292 | 75.0 | 0.612 | |
| Stat (mean+std+max) | 1.02M | 2.8 | 66.1 | 0.293 | 75.2 | 0.596 | |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Pairwise Accuracy (%) | Human Correlation | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | Backbone | ImplB p | VP2 p | PhyDEx p | VF2 p | |||||
| V-JEPA2 Surprise | V-JEPA2 ViT-H/16 | 33.3 | 46.3 | 49.0 | 45.5 | 0.065 | 0.065 | 0.121 | 0.133 | |
| PhyProbe | V-JEPA2 ViT-H/16 | 81.4 | 60.2 | 70.7 | 76.3 | 0.301 | 0.303 | 0.623 | 0.630 | |
| – | +48.1 | +13.9 | +21.7 | +30.8 | +0.236 | +0.238 | +0.502 | +0.497 | ||
| Minimum | # Test | Pairwise Accuracy (%) | |||
|---|---|---|---|---|---|
| Score Gap | Pairs | VideoScore | VideoPhy2-AutoEval | V-JEPA2 Surprise | PhyProbe |
| 1280 | 49.8 | 28.5 | 46.3 | 66.1 | |
| 628 | 52.0 | 32.8 | 44.6 | 71.2 | |
| 202 | 52.0 | 34.2 | 46.0 | 77.2 | |
| 35 | 37.1 | 45.7 | 48.6 | 74.3 | |
| Method | ImplB p | VP2 p | PhyDEx p | VF2 p |
| VP2-AutoEval | 28.0 | 28.5 | 30.5 | 36.3 |
| V-JEPA2 Surprise | 33.3 | 46.3 | 49.0 | 45.5 |
| VideoScore2 | 68.7 | 49.7 | 60.6 | 65.0 |
| PhyProbe | 98.0 | 66.1 | 98.9 | 75.2 |
| PhyProbe (w/o TRAVL) | 92.1 | 58.7 | 98.2 | 74.7 |
| (w/o TRAVL full) |
| Method | ImplB p | VP2 p | PhyDEx p | VF2 p |
|---|---|---|---|---|
| PhyProbe (linear) | 96.7 | 60.0 | 96.5 | 71.0 |
| PhyProbe | 98.0 | 66.1 | 98.9 | 75.2 |
| (linear proposed) |
| Subset | Linear | Proposed |
|---|---|---|
| ImplB p ( ) | 17.5% | 82.5% |
| VF2 p ( ) | 37.5% | 62.5% |
| VP2 p ( ) | 25.0% | 75.0% |
| Overall ( ) | 26.7% | 73.3% |
| Fleiss’ | 0.148 (4 raters) | |
| Unanimous agreement | 14/30 pairs (46.7%) | |
| Percentile | 5 | 10 | 20 | 30 | 40 |
|---|---|---|---|---|---|
| MonetBench | 69.8 | 71.4 | 72.7 | 71.8 | 72.4 |
| RewardBench Overall | 63.2 | 63.6 | 64.1 | 64.6 | 64.8 |
| RewardBench MQ | 57.9 | 59.4 | 62.3 | 65.0 | 67.9 |
| RewardBench VQ | 58.0 | 59.3 | 61.4 | 63.4 | 65.6 |
| Domain | Eval Dataset | Pairs | Type | Corres. | Desc. |
|---|---|---|---|---|---|
| Physical | ImplausiBench [ 36 ] | 150 | R–G | Img+Txt | Paired with real and physical violations |
| PhyDetEx [ 53 ] | 2K | R–G | None | Diverse real vs synthetic violations | |
| VideoPhy2-Test [ 6 ] | 1.2K | G–G | Txt | Paired with score differences | |
| VideoFeedback2-Test [ 19 ] | 2K | G–G | None | Paired with score differences | |
| Quality | MonetBench [ 55 ] | 1K | G–G | Txt | Human preference on video quality |
| Rapidata-I2V [ 45 , 46 ] | 295 | G–G | Img+Txt | Human preference on video quality |
| Evaluation Benchmark | Supervision overlap | Generation Prompt overlap | Generator overlap |
|---|---|---|---|
| Physical Domain | |||
| ImplB p (ImplausiBench) | 0/300 | Unavailable | Unavailable |
| VP2 p (VideoPhy2-test) | 0/2888 | 0/599 | 2/5 (1172/2943 samples) |
| PhyDEx p (PhyDetEx) | 0/398 | 0/247 | 2/3 (55/250 samples) |
| VF2 p (VideoFeedback2-test) | 0/500 | 500/500 | 25/25 (500/500 samples) |
| Quality Domain | |||
| Training | Evaluation | Supervision overlap | Generation Prompt overlap | Generator overlap |
|---|---|---|---|---|
| VideoPhy2-train | VideoPhy2-test | 0% | 0% | 43% |
| VideoFeedback2-train | VideoFeedback2-test | 0% | 100% | 100% |
| TRAVL | ImplausiBench | 0% | 0% | Unavailable |
| Field | Min | Q1 | Median | Q3 | P99 | Max |
|---|---|---|---|---|---|---|
| Mask area | 0.00011 | 0.014 | 0.039 | 0.094 | 0.446 | 0.860 |
| Center proximity | 0.00 | 0.34 | 0.56 | 0.79 | 1.00 | 1.00 |
| Mask presence ratio | 0.021 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Metric | Value |
|---|---|
| Heuristic vs. Human | |
| Annotator 1 agreement | 77.6% |
| Annotator 2 agreement | 74.1% |
| Annotator 3 agreement | 77.6% |
| Annotator 4 agreement | 65.9% |
| Human majority agreement | 80.6% (58/72) |
| Min | Q1 | Median | Q3 | IQR | Max |
| 0.012 | 0.24 | 0.31 | 0.38 | 0.14 | 0.699 |
| Eval Set | Model | Model tie rate | Acc (all pairs) | Acc (model non-tied) |
|---|---|---|---|---|
| VP2 p | VideoPhy2-AutoEval | 58.8% | 28.5% | 69.1% |
| VideoScore2 | 40.2% | 49.7% | 60.0% | |
| VF2 p | VideoPhy2-AutoEval | 46.0% | 36.3% | 67.2% |
| VideoScore2 | 38.1% | 65.0% | 79.6% | |
| ImplB p | VideoPhy2-AutoEval | 63.3% | 28.0% | 76.4% |
| VideoScore2 | 43.3% | 68.7% | 75.3% |