Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can often be myopic to physical dynamics, or fine-tuned evaluators trained on human annotations, which overfit to dataset-specific cues and fail to generalize. A key challenge is that existing supervision sources provide either relative ordering or absolute scores, but not both reliably and consistently across varied settings. To this end, we introduce PhyProbe, an evaluator that extracts features from a frozen pretrained spatio-temporal encoder and maps them to a scalar physical consistency violation score via a lightweight scoring head. PhyProbe is trained through a unified objective combining pairwise ranking, regression on noisy scalar annotations, and anchor-based calibration over a curated set of heterogeneous supervision sources. Experiments show that PhyProbe outperforms prior methods on most pairwise benchmarks spanning real-generated and generated-generated pairs under varying correspondence, with the largest gains in no-correspondence and generated-generated settings where existing fine-tuned evaluators degrade sharply. PhyProbe achieves strong correlation with human judgments, with close agreement between rank-based and linear metrics, indicating that scores are both well ordered and anchored to a stable [0, 1] scale. Further, despite being trained on supervision indicative of physical consistency, without explicit general-preference labels, PhyProbe also performs competitively on human preference benchmarks: consistent with the observation that physics violations are entangled with broader quality degradations.
Figures & tables
Figure 1 : Challenges in evaluating physical consistency. Generated videos can be visually realistic across the board, while their physical consistency varies from fully consistent to severely violated. Human perception and current VLM-based evaluators can be thresholded and inconsistent, making it difficult to capture the continuous spectrum of violation severity. This motivates our goal: learning a bounded, continuous measure of physical consistency violation.
Figure 2 : Overview of PhyProbe . Given a video, we extract a learned physical representation using a pretrained spatiotemporal visual encoder, followed by a lightweight scoring head to predict a continuous violation score s(v) . Training combines three complementary supervision signals: (i) pairwise ranking on video pairs (vi,vj) to model relative violation, (ii) regression on human-annotated scores ymag , and (iii) anchoring with calibration targets to enforce absolute scale (e.g., s(real)=0 , s(viol)=1 ). At inference time, the model operates on a single video without requiring pairwise comparison, producing a continuous estimate of physical violation severity.
Training Dataset
Pairs
Type
Correspondence
Rank
Mag
Cal
PAI-Bench [ 60 ]
42K
R-G
Image + Text
✓
✓
VideoPhy2 [ 6 ]
19K
G-G
Action / Text
✓
✓
VideoFeedback2 [ 19 ]
31K
G-G
None
✓
✓
BrokenVideos [ 28 ]
14K
G-G
None
✓
✓
ImpossibleVideos [ 2 ]
27K
R-G
None
✓
✓
TRAVL [ 36 ]
98K
R-G
None
✓
✓
Table 1: Training datasets used for PhyProbe . Each dataset contributes to different supervision signals: pairwise preference learning (rank), regression (mag), and anchor-based calibration (cal). During training, we sample data from different sources using dataset-level weights to balance their contributions, preventing large datasets from dominating the learning signal.
Pairwise Accuracy (%) ↑
Human Correlation ↑
Img+Txt
Text
None
Avg.
VP2
VF2
Avg.
Method
Backbone
ImplB p
VP2 p
PhyDEx p
VF2 p
ρ
r
ρ
r
ρ
r
(R–G)
(G–G)
(R–G)
(G–G)
1 VP2-AutoEval
VideoCon (7B)
28.0
28.5
30.5
36.3
32.1
0.359
0.362
0.235
0.240
0.298
0.302
2 V-JEPA Surprise
V-JEPA2-H/16 (600M)
33.3
46.3
49.0
45.5
46.6
0.065
0.065
0.121
0.133
0.093
0.099
3 VideoScore2
Qwen2.5-VL (7B)
68.7
49.7
60.6
65.0
59.9
0.197
0.160
0.462
0.430
0.336
0.301
Table 2 : Core results on physical video benchmarks. Left Columns: pairwise accuracy (%) disentangled by pair type (real (R) vs generated (G)) and correspondence type (image+text, text, none). Benchmark p indicates the benchmark is reordered in pairwise manner without ties. We additionally report a weighted average across all pairwise benchmarks based on the number of pairs in each benchmark. Right Columns: correlation with human judgments on VideoPhy2 (VP2) and VideoFeedback2 (VF2). We report Spearman’s rank correlation ( ρ ) and Pearson correlation ( r ) to assess both ranking consistency and linear alignment. We additionally report Fisher-z averaged correlations across datasets, weighting each benchmark equally. We further studied the prior method behaviors in Appendix D .
Figure 3 : Qualitative comparison of physical plausibility assessment across videos. Each column shows representative frames, along with human annotations and scores from PhyProbe and prior physical evaluators [ 6 , 19 ] . Arrows indicate whether the score is higher the better or vice versa.
Figure 4 : Qualitative results across correspondence levels. PhyProbe assigns low violation scores to videos with coherent object motion and interaction dynamics, while assigning higher scores to videos exhibiting implausible behaviors, effective across various correspondence level.
Method
MonetBench
Rapidata-I2V
VideoGen-RewardBench
w/ tie
w/o tie
w/o tie only
Overall
MQ
VQ
w/ tie
w/o tie
w/ tie
w/o tie
w/ tie
w/o tie
1 Random
33.2
49.9
49.7
33.2
49.7
33.2
49.5
33.3
49.8
2 V-JEPA Surprise
52.4
53.8
45.1
45.2
44.2
48.8
47.1
48.7
47.6
3 VideoScore2
42.3
41.1
45.8
57.4
58.9
54.3
60.6
55.3
60.1
4 VideoPhy-2-AutoEval
22.5
16.5
9.5
24.6
19.5
37.9
20.5
34.4
20.4
Table 3: PhyProbe aligns with broader human preference beyond physical consistency. Pairwise accuracy (%) on human preference (quality) benchmarks [ 55 , 45 , 46 , 29 ] under different tie-handling protocols. Despite being trained without explicit general-preference labels, PhyProbe achieves competitive performance across all datasets, categories (overall, MQ, VQ), and evaluation protocols. PhyProbe uses a model-side tie threshold at the 20th percentile.
Figure 5 : Distribution of evaluator predictions conditioned on human rating levels on VideoPhy2 (VP2) and VideoFeedback2 (VF2). From left to right: (a) PhyProbe on VP2, (b) VP2-AutoEval on VP2, (c) PhyProbe on VF2, and (d) VideoScore2 on VF2. Each violin shows the distribution of predicted scores for videos assigned the same human rating, with the red line indicating the mean. For visualization, PhyProbe violation scores are inverted so that higher values correspond to greater physical plausibility, matching the direction of the baseline evaluators. Compared with prior methods, PhyProbe characterizes physical inconsistencies with greater granularity than the five-point human scale, providing a fluid assessment of error intensity beyond rigid categorical labels.
Variant
Pairwise Accuracy (%)
Human Correlation
Score
None
Txt
Img+Txt
VP2
VF2
Range
PhyDEx p
VF2 p
VP2 p
ImplB p
ρ
r
ρ
r
(R–G)
(G–G)
(G–G)
(R–G)
PhyProbe ( −Lrank )
94.7
75.6
61.9
99.2
0.311
0.316
0.614
0.609
[−0.07,0.88]
PhyProbe ( −Lmag )
98.7
75.3
62.3
94.1
0.315
0.316
0.584
0.583
[−1.89,3.45]
PhyProbe ( −Lcal )
68.7
79.3
64.4
84.9
0.363
0.364
0.664
0.656
[−0.40,4.13]
Table 4: Importance of individual loss components in PhyProbe . We compare PhyProbe with different loss components using the same backbone (PE-Core-G14-448). The scoring range reflects the effective output range of each variant.
Pairwise Accuracy (%) ↑
Human Correlation ↑
Img+Txt
Text
None
VP2
VF2
PhyProbe
Params
ImplB p
VP2 p
PhyDEx p
VF2 p
ρ
r
ρ
r
Backbone
(R–G)
(G–G)
(R–G)
(G–G)
1 Wan2.2 VAE
300M
71.7
58.3
67.1
74.3
0.186
0.194
0.562
0.581
2 V-JEPA2-VIT-L/16
300M
76.1
60.6
88.8
74.2
0.228
0.206
0.564
0.580
3 V-JEPA2-VIT-H/16
600M
81.4
60.2
70.7
76.3
0.301
0.303
0.623
0.630
Table 5: Effect of backbone representations on PhyProbe . We train and evaluate PhyProbe with different pretrained video encoders, reporting pairwise accuracy (%) and correlation with human judgments. Pairwise accuracy is decomposed by correspondence type (image+text, text, none) and pair type (real (R) vs generated (G)). Benchmark p indicates the benchmark is reordered in pairwise manner without ties. Correlation metrics (Spearman ρ , Pearson r ) measure alignment with human judgments on VideoPhy2 (VP2) and VideoFeedback2 (VF2).
Head
Trainable
Input
Rel. (×)
VideoPhy2
VideoFeedback2
Params
Dim
Acc
ρ
Acc
ρ
Mean (MLP)
0.36M
[B,D]
1 ×
61.0
0.277
73.2
0.578
Attention Pooling
6.92M
[B,T,D]
19.2 ×
63.7
0.292
75.0
0.612
Stat (mean+std+max)
1.02M
[B,3D]
2.8 ×
66.1
0.293
75.2
0.596
Table 6: Ablation of projection head designs. We compare different spatiotemporal aggregation strategies in terms of performance and efficiency, under the same PECore backbone. Rel. denotes relative trainable parameter count normalized to the Mean (MLP) head.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Qualitative comparison of image-to-video (I2V) generations with varying levels of physical consistency. PhyProbe assigns low violation scores to videos with coherent object motion and interaction dynamics, while assigning higher scores to videos exhibiting implausible behaviors such as inconsistent motion, object deformation, or physical discontinuities. These examples illustrate that PhyProbe captures continuous variation in physical violation severity beyond binary plausible/implausible judgments.
Figure 7 : Qualitative comparison of text-to-video (T2V) generations under no semantic correspondence. Despite depicting entirely different scenes and prompts, PhyProbe is still able to form meaningful pairwise physical consistency comparisons by evaluating the severity of physical violations directly from video dynamics. Videos with coherent motion and interaction patterns receive lower violation scores, while videos exhibiting implausible motion, object deformation, or inconsistent dynamics receive higher scores.
Figure 8 : Failure cases of PhyProbe . In some cases, PhyProbe assigns elevated violation scores to real videos containing unusual or rarely observed motion patterns, suggesting limitations in the underlying representation space and training distribution. We also observe reduced sensitivity to violations that occur extremely briefly or instantaneously (e.g., sudden appearance, disappearance, or abrupt state changes), where the anomaly may not persist long enough to be strongly encoded in the aggregated video representation.
Pairwise Accuracy (%) ↑
Human Correlation ↑
Method
Backbone
ImplB p
VP2 p
PhyDEx p
VF2 p
ρVP2
rVP2
ρVF2
rVF2
V-JEPA2 Surprise
V-JEPA2 ViT-H/16
33.3
46.3
49.0
45.5
0.065
0.065
0.121
0.133
PhyProbe
V-JEPA2 ViT-H/16
81.4
60.2
70.7
76.3
0.301
0.303
0.623
0.630
Δ
–
+48.1
+13.9
+21.7
+30.8
+0.236
+0.238
+0.502
+0.497
Appendix
Table 7: We compare PhyProbe with V-JEPA2 suprise. Both methods use the identical V-JEPA2 ViT-H/16 backbone. PhyProbe consistently improves over the original V-JEPA2 Surprise baseline, indicating that the gains are not solely attributable to stronger backbone representations.
Minimum
# Test
Pairwise Accuracy (%)
Score Gap
Pairs
VideoScore
VideoPhy2-AutoEval
V-JEPA2 Surprise
PhyProbe
Δeval≥1
1280
49.8
28.5
46.3
66.1
Δeval≥2
628
52.0
32.8
44.6
71.2
Δeval≥3
202
52.0
34.2
46.0
77.2
Δeval≥4
35
37.1
45.7
48.6
74.3
Appendix
Table 8 : Pairwise accuracy (%) on the text-correspondence subset of VP2 p under different minimum ground-truth score-gap thresholds. Only video pairs generated from the same text prompt are considered, following the pairwise evaluation protocol; equal-score pairs are excluded. The same checkpoint for each method is evaluated across all rows, while only the evaluation subset is changed.
Method
ImplB p
VP2 p
PhyDEx p
VF2 p
VP2-AutoEval
28.0
28.5
30.5
36.3
V-JEPA2 Surprise
33.3
46.3
49.0
45.5
VideoScore2
68.7
49.7
60.6
65.0
PhyProbe
98.0
66.1
98.9
75.2
PhyProbe (w/o TRAVL)
92.1
58.7
98.2
74.7
Δ (w/o TRAVL − full)
−5.9
−7.4
−0.7
−0.5
Appendix
Table 9 : Leave-one-dataset-out (LODO) analysis. We retrain PhyProbe while excluding one training dataset (TRAVL, VideoPhy2, or VideoFeedback2) at a time and evaluate pairwise accuracy (%) across four physical video benchmarks. While performance degrades on the corresponding held-out benchmark and, in every case, on VP2 p ; it remains strong on the others, indicating partial transfer across sources alongside meaningful dependence on each dataset.
Method
ImplB p
VP2 p
PhyDEx p
VF2 p
PhyProbe (linear)
96.7
60.0
96.5
71.0
PhyProbe
98.0
66.1
98.9
75.2
Δ (linear − proposed)
−1.3
−6.1
−2.4
−4.2
Appendix
Table 10 : Effect of score normalization on PHYPROBE. We compare the proposed violation-flooring normalization against a simple linear normalization while keeping all other training settings identical. Results report pairwise accuracy (%) on four physical video benchmarks.
Subset
Linear
Proposed
ImplB p ( n=10 )
17.5%
82.5%
VF2 p ( n=10 )
37.5%
62.5%
VP2 p ( n=10 )
25.0%
75.0%
Overall ( n=30 )
26.7%
73.3%
Fleiss’ κ
0.148 (4 raters)
Unanimous agreement
14/30 pairs (46.7%)
Appendix
Table 11 : Blinded human preference evaluation of scalar target normalization. We compare the proposed violation-flooring normalization against a simple linear normalization using a blinded preference evaluation. Annotators were shown the predicted score differences from both models in random order and selected which better reflected the perceived difference in physical violation severity. Each of the n pairs was independently rated by four non-expert annotators; percentages are pooled over all annotator judgments. Given the small sample ( n=30 pairs) and inter-rater agreement of Fleiss’ κ=0.148 , we regard this as preliminary supporting evidence of a preference trend.
Figure 9 : Predicted score distributions for plausible and implausible videos. Score distributions produced by PhyProbe and VideoPhy2-AutoEval (VP2) on ImplB p (a,b) and VF2 p (c,d). For VF2 p , only samples with human scores of 1 and 5 are shown, corresponding to clearly implausible and plausible videos, respectively, to reduce ambiguity from intermediate ratings.
Percentile
5
10
20
30
40
MonetBench
69.8
71.4
72.7
71.8
72.4
RewardBench Overall
63.2
63.6
64.1
64.6
64.8
RewardBench MQ
57.9
59.4
62.3
65.0
67.9
RewardBench VQ
58.0
59.3
61.4
63.4
65.6
Appendix
Table 12: Sensitivity of PhyProbe “w/ tie” accuracy (%) to the model-side tie threshold, expressed as a percentile of absolute score differences on each benchmark. The 20th percentile is used for PhyProbe results in Table 3 . “w/o tie” columns compare raw scores without a threshold and are unaffected.
Figure 10 : Illustration of the supervision regime for physical plausibility evaluation. Human annotations primarily capture clearly observable violations (scores ≥0.5 ), while the region below the perceptual threshold remains largely unsupervised. Importantly, scores in the range [0,0.5) do not indicate the absence of violations, but rather subtle or weakly perceptible physical inconsistencies. PhyProbe produces differentiated scores in this region, but their interpretation remains unvalidated due to the absence of direct severity annotations. Figure 5 and Table 10 provide indirect evidence that finer resolution in the low-violation regime improves supervision.
Domain
Eval Dataset
Pairs
Type
Corres.
Desc.
Physical
ImplausiBench [ 36 ]
150
R–G
Img+Txt
Paired with real and physical violations
PhyDetEx [ 53 ]
2K
R–G
None
Diverse real vs synthetic violations
VideoPhy2-Test [ 6 ]
1.2K
G–G
Txt
Paired with score differences
VideoFeedback2-Test [ 19 ]
2K
G–G
None
Paired with score differences
Quality
MonetBench [ 55 ]
1K
G–G
Txt
Human preference on video quality
Rapidata-I2V [ 45 , 46 ]
295
G–G
Img+Txt
Human preference on video quality
Appendix
Table 13: Evaluation benchmarks used in this work, covering both physical plausibility and perceptual quality under diverse settings of correspondence and pair types. Physical benchmarks focus on consistency of dynamics, while quality benchmarks evaluate alignment with human preferences.
Figure 11 : Prompt template for pairwise physical consistency evaluation between two generated videos.
Figure 12 : Prompt template for scalar physical consistency scoring of a single video.
Evaluation Benchmark
Supervision overlap
Generation Prompt overlap
Generator overlap
Physical Domain
ImplB p (ImplausiBench)
0/300
Unavailable
Unavailable
VP2 p (VideoPhy2-test)
0/2888
0/599
2/5 (1172/2943 samples)
PhyDEx p (PhyDetEx)
0/398
0/247
2/3 (55/250 samples)
VF2 p (VideoFeedback2-test)
0/500
500/500
25/25 (500/500 samples)
Quality Domain
Appendix
Table 14 : Train/test overlap between each evaluation benchmark and the complete training corpus used by PhyProbe . Entries are reported as X/Y , where X denotes the number of unique evaluation video-label entities that also appear in the complete training corpus and Y denotes the total number of unique evaluation entities. Supervision overlap is computed over annotated supervision instances (e.g., preference labels, quality scores, or pairwise comparisons). Generation Prompt overlap is computed over unique conditioning inputs used to generate the evaluation videos (i.e., text prompts for text-to-video tasks and image–text input pairs for image-conditioned tasks). Generator overlap is computed over unique video generation models. Generator overlap does not imply supervision or train–test sample overlap.
Training
Evaluation
Supervision overlap
Generation Prompt overlap
Generator overlap
VideoPhy2-train
VideoPhy2-test
0%
0%
43%
VideoFeedback2-train
VideoFeedback2-test
0%
100%
100%
TRAVL
ImplausiBench
0%
0%
Unavailable
Appendix
Table 15 : Official train/test overlap for the benchmark datasets used in PhyProbe . Supervision overlap denotes identical annotated supervision instances shared between the official training and evaluation splits. Generation Prompt overlap denotes identical conditioning inputs (i.e., text prompts or image–text input pairs). Generator overlap denotes shared video generation models and is reported only when generator metadata is available. Note that VideoFeedback2’s official splits share prompts and generators; supervision is disjoint, but the test set does not constitute a prompt- or generator-level distribution shift.
Figure 13 : Human annotation distributions across different physical consistency evaluation datasets. Annotators consistently avoid extreme ratings and concentrate around intermediate scores, reflecting the coarse and thresholded nature of human perception toward physical inconsistencies. Anchor-based supervision serves as a mitigation strategy by providing stable reference points at the low- and high-violation ends of the spectrum, helping calibrate the absolute violation scale across heterogeneous datasets.
Field
Min
Q1
Median
Q3
P99
Max
Mask area
0.00011
0.014
0.039
0.094
0.446
0.860
Center proximity
0.00
0.34
0.56
0.79
1.00
1.00
Mask presence ratio
0.021
1.00
1.00
1.00
1.00
1.00
Appendix
Table 16: Distribution of per-video statistics in the BrokenVideos training set. Lower mask area, center proximity, and mask presence ratio indicate fewer localized anomalies. Annotation coverage ratio is reported only as a data quality statistic and is not used in the composite anomaly score.
Metric
Value
Heuristic vs. Human
Annotator 1 agreement
77.6%
Annotator 2 agreement
74.1%
Annotator 3 agreement
77.6%
Annotator 4 agreement
65.9%
Human majority agreement
80.6% (58/72)
Appendix
Table 17: Human validation of the proposed heuristic-derived cross-prompt BrokenVideos preference labels design, as the score is designed as 0.6⋅mask_area+0.3⋅center_proximity+0.1⋅mask_presence_ratio . Four independent non-expert annotators evaluated the same set of 85 sampled video pairs.
Min
Q1
Median
Q3
IQR
Max
0.012
0.24
0.31
0.38
0.14
0.699
Appendix
Table 18: Distribution of the composite anomaly score used for BrokenVideos pair construction. The score is computed as 0.6⋅mask_area+0.3⋅center_proximity+0.1⋅mask_presence_ratio . The interquartile range (IQR) characterizes the intrinsic spread of the heuristic score. The adopted cross-prompt threshold (0.4) is approximately 2.9× the IQR, making it a conservative threshold that retains only well-separated preference pairs.
Eval Set
Model
Model tie rate
Acc (all pairs)
Acc (model non-tied)
VP2 p
VideoPhy2-AutoEval
58.8%
28.5%
69.1%
VideoScore2
40.2%
49.7%
60.0%
VF2 p
VideoPhy2-AutoEval
46.0%
36.3%
67.2%
VideoScore2
38.1%
65.0%
79.6%
ImplB p
VideoPhy2-AutoEval
63.3%
28.0%
76.4%
VideoScore2
43.3%
68.7%
75.3%
Appendix
Table 19: Tie-aware pairwise accuracy of VideoPhy2-AutoEval and VideoScore2 across evaluation benchmarks. Both baselines output discrete physical-commonsense scores, so a pair receives a model tie whenever the two videos are assigned equal scores. Model tie rate is the fraction of pairs on which this occurs. Acc (all pairs) is computed over every pair, including human-labelled ties, with a model tie counted as correct only when the human label is also a tie; it corresponds to the “w/ tie” columns of Table 3. Acc (model non-tied) restricts evaluation to pairs on which the model produced a non-tied prediction, i.e., accuracy conditional on the model committing to a preference.
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Compounding this is a lack of reliable granular evaluation methods for localizing and specifying physical law violations in videos. We address this by introducing Physics Question Scene Graph (PQSG), a hierarchical question-based evaluation pipeline. PQSG evaluates generated videos by checking their faithfulness to a prompt across objects, actions, and adherence to physical laws using a graph-based hierarchy of questions generated by a vision-language model (VLM), guided by high-quality in-context examples. By representing questions as a graph, PQSG introduces logical dependencies within questions, ensuring that each query is contextually valid. Moreover, PQSG provides granular assessments of which qualities of the video violate physical plausibility constraints. We validate PQSG by creating FinePhyEval, a dataset with physics-based prompts and corresponding generated videos from diverse state-of-the-art video generation models (Sora 2, Veo 3, and Wan 2.1), with each video annotated across multiple categories by humans. Using FinePhyEval, we measure the correlation between PQSG's fine-grained scores and human judgments, showing higher overall correlations than prior work. We also find that PQSG ranks closed-source models higher than Wan 2.1 on physical realism. Lastly, we show that the annotations we provide in FinePhyEval can also be used for subtask evaluation: we benchmark two strong VLMs on generating and answering questions, finding that while models can create human-like questions, they still fall short of human performance in answering them.
Atin Pothiraj, Jaemin Cho, Yue Zhang +2
University of North Carolina at Chapel Hill, USA · AI2, USA · Johns Hopkins University, USA +1
Modern video generative models produce visually impressive results, yet frequently violate basic physical principles. We propose Proprio, a training-free framework that enables a frozen video generator to assess and improve the physical plausibility of its own outputs. Inspired by proprioception, the biological sense of one's own movement, Proprio treats the model's flow residual under controlled latent perturbations as a self-scoring signal. Samples that are better explained by the generator's learned dynamics induce smaller and more stable residuals. We aggregate this signal across timesteps and perturbations, focus it on motion-relevant regions with a dynamic spatiotemporal mask, and use it for best-of-N search, gradient-based self-refinement, or both. Across text-to-video and image-to-video benchmarks, Proprio consistently improves physical plausibility, outperforming VLM-based scoring, and external world-model baselines in several settings. With TurboWan2.2, Proprio improves Physics-IQ from 32.2 to 37.5 (+16.5%) and VideoPhy2-hard physical commonsense from 45.6 to 55.0 (+20.6%). Human evaluation further shows that raters prefer Proprio-selected or refined videos for physical plausibility in roughly two-thirds of comparisons. These results suggest that frozen video generators contain actionable internal signals for evaluating and improving the physical plausibility of their own outputs.
Mariam Hassan, Kaouther Messaoud, Wuyang Li +1
1 École Polytechnique Fédérale de Lausanne (EPFL) · Télécom Paris, IP Paris
We introduce reference-free measures for evaluating the physical consistency of generated videos, combining relative and absolute approaches to assess fidelity. Although tools like WorldGym or WorldEval enable robotic simulation via video generation, physical fidelity gaps often prevent these environments from accurately reproducing real-world task success rates of VLA models. Unlike existing evaluation methods, which require costly human voting (Elo) or unavailable ground-truth references (FVD), our approach utilizes DROID-SLAM and SEA-RAFT to quantify physical inconsistencies, motivated by WorldScore. Videos filtered using our relative consistency assessment show an improvement in task success rates of over 8%, effectively narrowing the simulation-to-reality gap. Furthermore, our absolute assessment enables spatio-temporal localization, providing visualization of when and where physical artifacts occur.