Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and failure causes. We further propose MotionInsight, a diagnostic evaluator that shifts assessment from implicit RGB-frame observation to explicit motion-space diagnosis. By constructing motion-aware representations, MotionInsight makes subtle motion deficiencies more observable. We also introduce motion-specific rewards during GRPO to transform observed motion into a diagnostic assessment. Experiments demonstrate that MotionInsight provides an effective basis for diagnosing object motion deficiencies, producing human-aligned scores along three dimensions and grounded explanations.
Figures & tables
Dim.
Description
OC
Stable identity, appearance, and structure during motion.
MC
Smooth and continuous trajectory without abrupt jumps or jitter.
PP
Motion consistent with forces and basic physical dynamics.
Table 1: Definitions of the three dimensions in object-centric motion fidelity assessment: object consistency (OC), motion continuity (MC), and physical plausibility (PP).
Model
Object Consistency ↑
Motion Continuity ↑
Physical Plausibility ↑
Open-source models
Wan 2.2 Wan et al. (2025)
3.420±0.441
3.634±0.530
2.355±0.261
LTX 2.1 HaCohen et al. (2024)
1.978±0.690
3.512±0.546
2.930±0.402
LongCat-Video Team et al. (2025)
1.335±0.214
2.242±0.699
1.821±0.582
Cosmos-Predict 2.5 Ali et al. (2025)
2.088±0.649
3.109±0.679
2.156±0.472
HunyuanVideo-1.5 Kong et al. (2024)
3.223±0.504
3.939±0.672
2.370±0.274
Table 2: Comparison of video generation models and real videos on object consistency, motion continuity, and physical plausibility, evaluated on the 166 challenging prompts in VidMotion-Test. Values are reported as mean ± standard deviation.
Figure 2: Overview of MotionInsight. Given an input video and a target prompt, we uniformly sample RGB frames as the standard visual input. Then, we extract motion features from the frames using ViPE and CoTracker3, and aggregate them into a motion-aware representation, which is fed into the VLM.
Figure 3: Qualitative result of MotionInsight. The reasoning texts are presented below the video frames, and the scores for the three dimensions are visualized in the radar chart at the top right.
Object Consistency
Motion Continuity
Physical Plausibility
Method
SRCC ↑
PLCC ↑
KRCC ↑
SRCC ↑
PLCC ↑
KRCC ↑
SRCC ↑
PLCC ↑
KRCC ↑
Independent Human
0.732
0.774
0.671
0.614
0.684
0.580
0.766
0.804
0.692
GPT-5.4
0.370
0.400
0.322
0.276
0.303
0.243
0.246
0.257
0.203
Gemini 3.1 Pro
0.333
0.318
0.281
0.218
0.217
0.193
0.166
0.162
0.144
Qwen-3-VL-8B
0.074
0.092
0.065
0.055
0.062
0.051
-0.009
0.020
-0.008
Qwen-3-VL-8B-FT
0.286
0.301
0.242
0.213
0.238
0.190
0.181
0.205
0.154
Table 3: Correlation with human annotations across three dimensions. Baselines include GPT-5.4 Hurst et al. (2024) , Gemini 3.1 Pro Team et al. (2024) , Qwen-3-VL-8B Bai et al. (2025) , Qwen-3-VL-8B-FT (supervised fine-tuning on VidMotion-Train), VideoPhy2 Bansal et al. (2025) , VMBench Ling et al. (2025) , and WorldModelBench Li et al. (2025) . Independent Human denotes a single volunteer evaluator who did not participate in dataset annotation or calibration, while the benchmark labels are aggregated from three calibrated annotators using MOS.
Method
Jaccard ↑
Prec. ↑
Rec. ↑
GPT-5.4
0.15
0.29
0.19
Gemini 3.1 Pro
0.12
0.24
0.16
Qwen-3-VL-8B
0.09
0.20
0.12
Ours
0.58
0.74
0.63
Table 4: Failure-cause grounding of diagnostic reasoning. We compare GPT-5.4 Hurst et al. (2024) , Gemini 3.1 Pro Team et al. (2024) , Qwen-3-VL-8B Bai et al. (2025) , and our MotionInsight. We parse each model’s reasoning into 12 predefined failure causes and compare the parsed causes with human annotations using Jaccard similarity, precision, and recall.
Figure 4: Sensitivity to localized motion failures. We divide each video into K temporal clips and compute the gap between the minimum clip-level score and the human full-video score. MotionInsight keeps a consistently small gap across K , indicating better sensitivity.
Method
#Fr.
OC ↑
MC ↑
PP ↑
Uniform Sampling
16
0.510
0.455
0.445
2 × Uniform Sampling
32
0.435
0.382
0.421
4 × Uniform Sampling
64
0.406
0.334
0.397
Trajectory Overlay
16
0.475
0.438
0.430
Motion-Aware Rep.
16
0.761
0.633
0.727
Table 5: Ablation on video perception strategies under the same GRPO training, reported with SRCC. Rep. denotes representations. #Fr. refers to the number of sampled RGB frames, OC: Object Consistency, MC: Motion Continuity, and PP: Physical Plausibility.
Setting
OC ↑
MC ↑
PP ↑
Semantic Alignment Only
0.253
0.183
0.224
+ Scoring Reward
0.703
0.588
0.670
+ Failure-cause Reward
0.761
0.633
0.727
Table 6: Ablation on the GRPO design of MotionInsight, reported with SRCC. OC: Object Consistency, MC: Motion Continuity, and PP: Physical Plausibility.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Method
OC ↑
MC ↑
PP ↑
Motion-Aware Rep.
0.761
0.633
0.727
w/ Temporal Shuffle
0.161
0.109
0.113
w/o Camera Motion
0.697
0.580
0.683
Appendix
Table 7: Ablation on the design of motion-aware representations for GRPO, reported with SRCC. Rep. denotes representations. w/ Temporal Shuffle randomly permutes the temporal order of the motion representations, while w/o Camera Motion removes camera-motion tokens and keeps only object-motion features. OC: Object Consistency, MC: Motion Continuity, and PP: Physical Plausibility.
Figure 5: Distribution of target objects in VidMotion. The objects are grouped into six coarse categories: Human, Animal, Vehicle, Toy and Ball, Tool, and Other. The distribution highlights the diversity of object-centric motion scenarios covered by the dataset.
Dimension
Krippendorff’s α
Object Consistency
0.7981
Motion Continuity
0.6846
Physical Plausibility
0.7432
Appendix
Table 8: Inter-annotator reliability for the three scalar motion-quality dimensions in VidMotion. Krippendorff’s α is used to measure agreement among annotators, with higher values indicating stronger consistency.
Method
Motion Fidelity
Overall Quality
over Wan 2.1
92.86%
86.67%
over VideoPhy2-DPO
87.56%
87.50%
Appendix
Table 9: 2AFC human study results for DPO-based video generation. We compare the model trained with MotionInsight-derived preference pairs against the original Wan2.1 and the VideoPhy2-DPO Bansal et al. (2025) baseline.
Figure 6: Qualitative comparison of DPO-based video generation.
Figure 7: Interface of the annotation website.
Figure 8: Distribution of annotated issue categories. The 12 cause candidates are grouped into three dimensions: Object Consistency, Motion Continuity, and Physical Plausibility.
Figure 9: Additional qualitative results of MotionInsight. For each example, we show the sampled video frames, the model’s reasoning for the designated object, and the predicted scores across the three motion-quality dimensions.
Figure 10: Overview of the VidMotion construction and annotation pipeline. We collect real videos with prominent object motion, extract captions and target objects, generate corresponding videos using nine open-source and proprietary models, and collect expert annotations for both real and generated videos, including three-dimensional scores, failure causes, and confidence levels.
Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method systematically measures how faithfully these systems render human bodies and motion dynamics. In this paper, we present HumanScore, a systematic framework to evaluate the quality of human motions in AI-generated videos. HumanScore defines six interpretable metrics spanning kinematic plausibility, temporal stability, and biomechanical consistency, enabling fine-grained diagnosis beyond visual realism alone. Through carefully designed prompts, we elicit a diverse set of movements at varying intensities and evaluate videos generated by thirteen state-of-the-art models. Our analysis reveals consistent gaps between perceptual plausibility and motion biomechanical fidelity, identifies recurrent failure modes (e.g., temporal jitter, anatomically implausible poses, and motion drift), and produces robust model rankings from quantitative and physically meaningful criteria.
Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Models (MLLMs) are strongly appearance-biased, lack temporal understanding, and thus struggle to discern intricate motion dynamics and anatomical implausibilities in generated videos. We tackle this gap by introducing a novel evaluation metric derived from a learned latent space of real-world human actions. Our method first captures the nuances, constraints, and temporal smoothness of real-world motion by fusing appearance-agnostic human skeletal geometry features with appearance-based features. We posit that this combined feature space provides a robust representation of action plausibility. Given a generated video, our metric quantifies its action quality by measuring the distance between its underlying representations and this learned real-world action distribution. For rigorous validation, we develop a new multi-faceted benchmark specifically designed to probe temporally challenging aspects of human action fidelity. Through extensive experiments, we show that our metric achieves substantial improvement of more than 68% compared to existing state-of-the-art methods on our benchmark, performs competitively on established external benchmarks, and has a stronger correlation with human perception. Our in-depth analysis reveals critical limitations in current video generative models and establishes a new standard for advanced research in video generation.
Xavier Thomas, Youngsun Lim, Ananya Srinivasan +2
Boston University · Belmont High School · Canyon Crest Academy +1
Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and motion remains challenging. Most existing video evaluation pipelines rely heavily on human judgment or learned graders, which can be subjective and weakly diagnostic for geometric failures. We introduce PDI-Bench (Perspective Distortion Index), a quantitative framework for auditing geometric coherence in generated videos. Given a generated clip, we obtain object-centric observations via segmentation and point tracking (e.g., SAM 2, MegaSaM, and CoTracker3), lift them to 3D world-space coordinates via monocular reconstruction, and compute a set of projective-geometry residuals capturing three failure dimensions: scale-depth alignment, 3D motion consistency, and 3D structural rigidity. To support systematic evaluation, we build PDI-Dataset, covering diverse scenarios designed to stress these geometric constraints. Across state-of-the-art video generators, PDI reveals consistent geometry-specific failure modes that are not captured by common perceptual metrics, and provides a diagnostic signal for progress toward physically grounded video generation and physical world model. Our code and dataset can be found at https://pdi-bench.github.io/.
Jiaxin Wu, Yihao Pi, Yinling Zhang +2
1Tsinghua University - IEI Lab · 2UW-Madison · 3Adobe Research