cs.CLSep 29, 2026

Momentum-Coupled Rubric Adaptation for Detailed Image Captioning

Authors: Zhenwen Ji, Lei Jin, Shanyong Wang, Jiaming Lu, Chengqiang Lu, Yi Wu, Yao Hu, Lizhen Cui, +1 more

Organizations: Xiaohongshu Inc. · the Joint SDU-NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University

Abstract

Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision--language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

    Jul 6, 2026Mingqi Gao, Hongyuan Dong, Yifei Chen +6Video CaptioningCategory-Aware Atomic Claims

  2. RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

    Nov 25, 2025Xuelu Feng, Yunsheng Li, Ziyu Wan +4Modern Text-To-Image ModelsText-To-Image