cs.CVSep 29, 2026

Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning

Authors: Yanan Wang, Tingsong Li, Kaixun Jiang, Chongyang Zhong, Chenwei Xoe, Zhaohe Liao

Organizations: Zhejiang University · Alibaba Token Hub, Alibaba Group · University of Science and Technology of China · Fudan University · Shanghai Jiao Tong University

Abstract

Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and shot contexts. Our Graded Physical Alignment score (GPA) awards credit for correct content and deducts points for incorrect or fabricated actions, making these errors explicit in the aggregate score. Building on this rubric, we propose Graded Margin Direct Preference Optimization (GM-DPO), which assigns stronger preference margins and greater training weight to more severe action errors. Across three VLM backbones, GM-DPO achieves the highest substantive-action and GPA scores among the evaluated preference objectives, improving GPA over DPO by 2.02-3.40 points. On Qwen3-8B, it reduces the weighted hallucination rate by 21.3% relative to DPO. These gains accompany sustained long-form output, improved shot structure, and competitive performance on three additional multimodal benchmarks.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

    Jun 3, 2026Yong Cao, Chuqiao Li, Xianghui Xie +2Visual Question Answering BenchmarksHuman Motion

  2. Building a Precise Video Language with Human-AI Oversight

    Apr 22, 2026Zhiqiu Lin, Chancharik Mitra, Siyuan Cen +13Video-Language ModelsCinematic Ideal

  3. MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

    Jun 28, 2026Weisong Liu, Haochen Wang, Kuan Gao +8Video CaptioningMotion