Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and failure causes. We further propose MotionInsight, a diagnostic evaluator that shifts assessment from implicit RGB-frame observation to explicit motion-space diagnosis. By constructing motion-aware representations, MotionInsight makes subtle motion deficiencies more observable. We also introduce motion-specific rewards during GRPO to transform observed motion into a diagnostic assessment. Experiments demonstrate that MotionInsight provides an effective basis for diagnosing object motion deficiencies, producing human-aligned scores along three dimensions and grounded explanations.
Figures & tables
Dim.
Description
OC
Stable identity, appearance, and structure during motion.
MC
Smooth and continuous trajectory without abrupt jumps or jitter.
PP
Motion consistent with forces and basic physical dynamics.
Table 1: Definitions of the three dimensions in object-centric motion fidelity assessment: object consistency (OC), motion continuity (MC), and physical plausibility (PP).
Model
Object Consistency ↑
Motion Continuity ↑
Physical Plausibility ↑
Open-source models
Wan 2.2 Wan et al. (2025)
3.420±0.441
3.634±0.530
2.355±0.261
LTX 2.1 HaCohen et al. (2024)
1.978±0.690
3.512±0.546
2.930±0.402
LongCat-Video Team et al. (2025)
1.335±0.214
2.242±0.699
1.821±0.582
Cosmos-Predict 2.5 Ali et al. (2025)
2.088±0.649
3.109±0.679
2.156±0.472
HunyuanVideo-1.5 Kong et al. (2024)
3.223±0.504
3.939±0.672
2.370±0.274
Table 2: Comparison of video generation models and real videos on object consistency, motion continuity, and physical plausibility, evaluated on the 166 challenging prompts in VidMotion-Test. Values are reported as mean ± standard deviation.
Figure 2: Overview of MotionInsight. Given an input video and a target prompt, we uniformly sample RGB frames as the standard visual input. Then, we extract motion features from the frames using ViPE and CoTracker3, and aggregate them into a motion-aware representation, which is fed into the VLM.
Figure 3: Qualitative result of MotionInsight. The reasoning texts are presented below the video frames, and the scores for the three dimensions are visualized in the radar chart at the top right.
Object Consistency
Motion Continuity
Physical Plausibility
Method
SRCC ↑
PLCC ↑
KRCC ↑
SRCC ↑
PLCC ↑
KRCC ↑
SRCC ↑
PLCC ↑
KRCC ↑
Independent Human
0.732
0.774
0.671
0.614
0.684
0.580
0.766
0.804
0.692
GPT-5.4
0.370
0.400
0.322
0.276
0.303
0.243
0.246
0.257
0.203
Gemini 3.1 Pro
0.333
0.318
0.281
0.218
0.217
0.193
0.166
0.162
0.144
Qwen-3-VL-8B
0.074
0.092
0.065
0.055
0.062
0.051
-0.009
0.020
-0.008
Qwen-3-VL-8B-FT
0.286
0.301
0.242
0.213
0.238
0.190
0.181
0.205
0.154
Table 3: Correlation with human annotations across three dimensions. Baselines include GPT-5.4 Hurst et al. (2024) , Gemini 3.1 Pro Team et al. (2024) , Qwen-3-VL-8B Bai et al. (2025) , Qwen-3-VL-8B-FT (supervised fine-tuning on VidMotion-Train), VideoPhy2 Bansal et al. (2025) , VMBench Ling et al. (2025) , and WorldModelBench Li et al. (2025) . Independent Human denotes a single volunteer evaluator who did not participate in dataset annotation or calibration, while the benchmark labels are aggregated from three calibrated annotators using MOS.
Method
Jaccard ↑
Prec. ↑
Rec. ↑
GPT-5.4
0.15
0.29
0.19
Gemini 3.1 Pro
0.12
0.24
0.16
Qwen-3-VL-8B
0.09
0.20
0.12
Ours
0.58
0.74
0.63
Table 4: Failure-cause grounding of diagnostic reasoning. We compare GPT-5.4 Hurst et al. (2024) , Gemini 3.1 Pro Team et al. (2024) , Qwen-3-VL-8B Bai et al. (2025) , and our MotionInsight. We parse each model’s reasoning into 12 predefined failure causes and compare the parsed causes with human annotations using Jaccard similarity, precision, and recall.
Figure 4: Sensitivity to localized motion failures. We divide each video into K temporal clips and compute the gap between the minimum clip-level score and the human full-video score. MotionInsight keeps a consistently small gap across K , indicating better sensitivity.
Method
#Fr.
OC ↑
MC ↑
PP ↑
Uniform Sampling
16
0.510
0.455
0.445
2 × Uniform Sampling
32
0.435
0.382
0.421
4 × Uniform Sampling
64
0.406
0.334
0.397
Trajectory Overlay
16
0.475
0.438
0.430
Motion-Aware Rep.
16
0.761
0.633
0.727
Table 5: Ablation on video perception strategies under the same GRPO training, reported with SRCC. Rep. denotes representations. #Fr. refers to the number of sampled RGB frames, OC: Object Consistency, MC: Motion Continuity, and PP: Physical Plausibility.
Setting
OC ↑
MC ↑
PP ↑
Semantic Alignment Only
0.253
0.183
0.224
+ Scoring Reward
0.703
0.588
0.670
+ Failure-cause Reward
0.761
0.633
0.727
Table 6: Ablation on the GRPO design of MotionInsight, reported with SRCC. OC: Object Consistency, MC: Motion Continuity, and PP: Physical Plausibility.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Method
OC ↑
MC ↑
PP ↑
Motion-Aware Rep.
0.761
0.633
0.727
w/ Temporal Shuffle
0.161
0.109
0.113
w/o Camera Motion
0.697
0.580
0.683
Appendix
Table 7: Ablation on the design of motion-aware representations for GRPO, reported with SRCC. Rep. denotes representations. w/ Temporal Shuffle randomly permutes the temporal order of the motion representations, while w/o Camera Motion removes camera-motion tokens and keeps only object-motion features. OC: Object Consistency, MC: Motion Continuity, and PP: Physical Plausibility.
Figure 5: Distribution of target objects in VidMotion. The objects are grouped into six coarse categories: Human, Animal, Vehicle, Toy and Ball, Tool, and Other. The distribution highlights the diversity of object-centric motion scenarios covered by the dataset.
Dimension
Krippendorff’s α
Object Consistency
0.7981
Motion Continuity
0.6846
Physical Plausibility
0.7432
Appendix
Table 8: Inter-annotator reliability for the three scalar motion-quality dimensions in VidMotion. Krippendorff’s α is used to measure agreement among annotators, with higher values indicating stronger consistency.
Method
Motion Fidelity
Overall Quality
over Wan 2.1
92.86%
86.67%
over VideoPhy2-DPO
87.56%
87.50%
Appendix
Table 9: 2AFC human study results for DPO-based video generation. We compare the model trained with MotionInsight-derived preference pairs against the original Wan2.1 and the VideoPhy2-DPO Bansal et al. (2025) baseline.
Figure 6: Qualitative comparison of DPO-based video generation.
Figure 7: Interface of the annotation website.
Figure 8: Distribution of annotated issue categories. The 12 cause candidates are grouped into three dimensions: Object Consistency, Motion Continuity, and Physical Plausibility.
Figure 9: Additional qualitative results of MotionInsight. For each example, we show the sampled video frames, the model’s reasoning for the designated object, and the predicted scores across the three motion-quality dimensions.
Figure 10: Overview of the VidMotion construction and annotation pipeline. We collect real videos with prominent object motion, extract captions and target objects, generate corresponding videos using nine open-source and proprietary models, and collect expert annotations for both real and generated videos, including three-dimensional scores, failure causes, and confidence levels.