Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning
Organizations: Zhejiang University · Alibaba Token Hub, Alibaba Group · University of Science and Technology of China · Fudan University · Shanghai Jiao Tong University
Abstract
Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and shot contexts. Our Graded Physical Alignment score (GPA) awards credit for correct content and deducts points for incorrect or fabricated actions, making these errors explicit in the aggregate score. Building on this rubric, we propose Graded Margin Direct Preference Optimization (GM-DPO), which assigns stronger preference margins and greater training weight to more severe action errors. Across three VLM backbones, GM-DPO achieves the highest substantive-action and GPA scores among the evaluated preference objectives, improving GPA over DPO by 2.02-3.40 points. On Qwen3-8B, it reduces the weighted hallucination rate by 21.3% relative to DPO. These gains accompany sustained long-form output, improved shot structure, and competitive performance on three additional multimodal benchmarks.
Figures & tables
| Model | Motion Accuracy | Motion Sequence | Output Statistics | |||||||
| Subs. | Non-subs. | GPA | Len | SMR (%) | ||||||
| Proprietary Frontier APIs | ||||||||||
| Gemini-3.1-Pro | 17.80 | 42.21 | 27.56 | 34.27 | 26.53 | 28.18 | 30.04 | 28.78 | 855.7 | 50.27 |
| Gemini-3.5-Flash | 16.70 | 43.61 | 27.46 | 33.98 | 27.03 | 26.88 | 28.77 | 27.86 | 709.4 | 56.22 |
| Seed-2.1-Pro | 32.24 | 44.38 | 37.10 | 44.92 | 47.00 | 44.14 | 47.31 | 46.30 | 782.3 | 68.31 |
| Kimi-k2.6 | 18.47 | 42.13 | 27.93 | 36.04 | 30.35 | 30.26 | 32.94 | 31.62 | 683.4 | 53.57 |
| Model Variant | Motion Accuracy | Motion Sequence | Output Statistics | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Subs. | Non-subs. | GPA | Len | SMR (%) | ||||||
| Qwen3-8B-SFT | 6.92 | 42.31 | 21.08 | 29.88 | 13.30 | 15.54 | 15.81 | 15.23 | 1634 | 89.7 |
| Qwen2.5-7B-Base | 3.02 | 23.17 | 11.08 | 14.26 | 8.86 | 9.88 | 10.38 | 9.93 | 1131 | 49.0 |
| Qwen2.5-7B-DPO | 3.92 | 24.22 | 12.04 | 25.05 | 15.13 | 15.96 | 18.39 | 17.01 | 1420 | 94.0 |
| Qwen2.5-7B-cDPO | 3.05 | 25.80 | 12.15 | 25.20 | 15.25 | 14.55 | 17.65 | 16.24 | 1303 | 96.0 |
| Qwen2.5-7B-IPO | 2.94 | 27.14 | 12.62 | 26.40 | 17.83 | 18.42 | 19.22 | 18.70 | 1521 | 98.7 |
| Model | GPA | ||||
|---|---|---|---|---|---|
| Qwen2.5 | ✗ | ✗ | 12.04 | 25.05 | 17.01 |
| ✓ | ✗ | 12.29 | 25.00 | 16.52 | |
| ✗ | ✓ | 13.20 | 25.44 | 17.73 | |
| ✓ | ✓ | 14.06 | 25.82 | 18.46 | |
| Qwen3 | ✗ | ✗ | 19.66 | 28.86 | 19.78 |
| ✓ | ✗ | 19.93 | 29.42 | 19.88 |
| Method Variant | V-MME | POPE | V-GPT | |
|---|---|---|---|---|
| Acc | Acc | F1 | Score | |
| Q2.5-DPO | 60.06 | 87.43 | 86.06 | 2.50 |
| Q2.5-cDPO | 59.58 | 87.46 | 86.06 | 2.50 |
| Q2.5-IPO | 60.45 | 87.44 | 86.09 | 2.46 |
| Q2.5-SimPO | 60.26 | 87.41 | 86.02 | 2.48 |
| Q2.5-GM-DPO | 60.62 | 87.56 | 86.15 | 2.51 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Preference Loss | Hyperparameters |
|---|---|---|
| DPO ( Rafailov et al., 2023 ) | ||
| cDPO ( Mitchell, 2023 ) | ||
| IPO ( Azar et al., 2024 ) | ||
| SimPO ( Meng et al., 2024 ) | ||
| GM-DPO (Ours) |
| Method | Motion Accuracy | Motion Sequence | ||||||
|---|---|---|---|---|---|---|---|---|
| Subs. | Non-subs. | GPA | ||||||
| Qwen2.5-7B | ||||||||
| DPO | ||||||||
| cDPO | ||||||||
| IPO | ||||||||
| SimPO | ||||||||
| Backbone | GPA Gain | Mean | 95% CI | ||
|---|---|---|---|---|---|
| Run 1 | Run 2 | Run 3 | |||
| Qwen2.5-7B | |||||
| Qwen3-8B | |||||
| LLaVA-2-8B | |||||
| Model | Weight | Margin | Motion Accuracy (Penalty-Aware) | Motion Sequence (Non-Penalty) | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Subs. | Non-subs. | GPA | ||||||||
| Qwen2.5-7B | ✗ | ✗ | 3.92 | 24.22 | 12.04 | 25.05 | 15.13 | 15.96 | 18.39 | 17.01 |
| ✓ | ✗ | 4.19 | 24.44 | 12.29 | 25.00 | 15.86 | 14.61 | 17.92 | 16.52 | |
| ✗ | ✓ | 5.66 | 24.51 | 13.20 | 25.44 | 16.16 | 15.89 | 19.47 | 17.73 | |
| ✓ | ✓ | 6.14 | 25.95 | 14.06 | 25.82 | 17.05 | 17.31 | 19.72 | 18.46 | |
| Qwen3-8B | ✗ | ✗ | 7.16 | 38.40 | 19.66 | 28.86 | 17.69 | 19.94 | 20.52 | 19.78 |
| Model Variant | Short | Medium | Long | Overall |
|---|---|---|---|---|
| Qwen2.5-7B-Base | 68.89 | 59.11 | 50.11 | 59.37 |
| Qwen2.5-7B-DPO | 69.11 | 60.18 | 50.89 | 60.06 |
| Qwen2.5-7B-cDPO | 68.89 | 59.53 | 50.33 | 59.58 |
| Qwen2.5-7B-IPO | 69.56 | 60.18 | 51.61 | 60.45 |
| Qwen2.5-7B-SimPO | 69.40 | 60.05 | 51.34 | 60.26 |
| Qwen2.5-7B-GM-DPO (Ours) | 69.75 | 60.17 | 51.93 | 60.62 |
| Model Variant | Adversarial | Popular | Random | Overall Acc | Yes% | F1 |
|---|---|---|---|---|---|---|
| Qwen2.5-7B-Base | 86.50 | 87.37 | 88.23 | 87.37 | 40.03 | 85.97 |
| Qwen2.5-7B-DPO | 86.53 | 87.43 | 88.33 | 87.43 | 40.14 | 86.06 |
| Qwen2.5-7B-cDPO | 86.57 | 87.47 | 88.33 | 87.46 | 40.12 | 86.06 |
| Qwen2.5-7B-IPO | 86.53 | 87.40 | 88.40 | 87.44 | 40.29 | 86.09 |
| Qwen2.5-7B-SimPO | 86.50 | 87.38 | 88.36 | 87.41 | 40.10 | 86.02 |
| Qwen2.5-7B-GM-DPO (Ours) | 86.67 | 87.53 | 88.47 | 87.56 | 40.12 | 86.15 |
| Generic Understanding | Temporal | Consistency | Overall | |||
|---|---|---|---|---|---|---|
| Model Variant | Correctness | Detail | Context | Reasoning | ||
| Qwen2.5-7B-Base | 2.33 | 2.23 | 2.63 | 2.16 | 2.91 | 2.45 |
| Qwen2.5-7B-DPO | 2.32 | 2.33 | 2.64 | 2.22 | 2.98 | 2.50 |
| Qwen2.5-7B-cDPO | 2.31 | 2.35 | 2.63 | 2.23 | 2.98 | 2.50 |
| Qwen2.5-7B-IPO | 2.30 | 2.29 | 2.58 | 2.19 | 2.92 | 2.46 |
| Qwen2.5-7B-SimPO | 2.31 | 2.32 | 2.60 | 2.23 | 2.96 | 2.48 |