Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision--language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.
Figures & tables
Figure 1: Comparison of heterogeneous and shared model distributions in rubric-based caption optimization.
Figure 2: Overview of MoCo Rubric. Stage I trains the three roles through shared multi-role SFT. In Stage II, current Policy candidates inform the rubric, the momentum Generator and Judge assign rewards, and only the Policy receives RL gradients. The shared rubric model is updated through EMA at synchronization boundaries.
Method
PixMo-Cap
DenseFusion
CapArena
CompreCap
DOCCI
Average
Captioning Baselines
ShareGPT4V-7B
2.02
1.80
2.67
5.21
2.41
2.82
RICO-Flash-7B
8.89
5.41
9.00
9.87
9.04
8.44
OmniCaptioner-8B
10.91
10.62
12.17
10.05
11.24
11.00
JoyCaption-8B
11.72
4.81
12.67
25.85
14.46
13.90
MetaCaptioner-8B
22.63
23.25
19.00
32.50
18.47
23.17
Table 1: Pairwise win rates (%) against Qwen3-VL-8B across five captioning benchmarks. Average denotes the benchmark mean. Best results are bold; second-best results are underlined.
Method
CaptionQA
BLINK
TextVQA
DocVQA
ChartQA
Average
Base
80.41
47.05
58.00
69.18
61.28
63.18
RubiCap
83.28
49.05
56.85
77.90
64.44
66.31
Qwen3-VL-32B
82.92
50.94
56.69
75.06
62.16
65.55
EvoLM
82.00
49.00
57.08
77.21
62.44
65.55
SFT (shared)
81.05
47.69
56.44
76.13
63.16
64.89
Ours
84.24
50.10
58.33
79.05
64.12
67.17
Table 2: Caption-based question answering across five benchmarks. Average is the mean across benchmarks. Bold and underlined values indicate the best and second-best results, respectively.
ID
SFT initialization
Rubrics
Rubric-role update
PixMo-Cap
DenseFusion
Average
M1
Policy only
—
—
57.37
56.91
57.14
M2
Shared across roles
—
—
55.15
53.31
54.23
M3
Separate per role
Offline, fixed
No momentum update
70.74
66.00
68.37
M4
Separate per role
Online
No momentum update
71.31
67.13
69.22
M5
Separate per role
Online
EMA ( m=0.99 , H=1 )
65.06
61.80
63.43
M6
Shared across roles
Online
No momentum update
68.48
64.73
66.61
Table 3: Component ablation on PixMo-Cap and DenseFusion, reported as win rates (%). Bold and underlined values indicate the best and second-best available results, respectively.
Figure 6
Figure 4: Left: Momentum coefficient analysis at H=1 , with tested coefficients equally spaced and the frozen reference ( m=1 ) shown separately. Middle: Synchronization interval analysis at m=0.99 . Markers indicate measured results; lines connect adjacent tested settings. Right: Policy–momentum tracking at H=1 . Frozen and full-sync curves are boundary references.
m
H=1
H=5
H=20
0.95
69.58
68.18
69.29
0.99
72.14
71.79
70.18
0.999
68.28
66.67
69.48
Table 5: Equal-weight mean win rate (%) across PixMo-Cap and DenseFusion in the joint m×H sweep.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
PixMo-Cap
DenseFusion
CapArena
CompreCap
DOCCI
Average
Captioning Baselines
ShareGPT4V-7B
1.00
0.20
0.50
1.25
1.00
0.79
RICO-Flash-7B
5.02
4.02
5.34
8.06
4.20
5.33
OmniCaptioner-8B
5.62
3.41
6.34
7.89
9.80
6.61
JoyCaption-8B
11.65
6.43
11.52
31.18
13.60
14.87
MetaCaptioner-8B
14.06
10.64
13.52
20.97
14.80
14.80
Appendix
Table 6: Pairwise win rates (%) against Qwen3-VL-8B with GPT-5.6 Sol as the evaluator. Average denotes the benchmark mean. Best results are bold; second-best results are underlined.
LLM judge
Exact agreement
Gemini 3.1 Pro
429/499 (85.97%)
GPT-5.6 Sol
412/497 (82.90%)
Appendix
Table 7: Agreement between LLM judges and final human preferences on Qwen3-8B versus GPT-5.6 caption pairs.
Figure 5: GRPO training dynamics of RubiCap and MoCo Rubric. The upper panels show mean training reward; the lower panels show the fraction of rollout groups with zero within-group reward standard deviation. Light traces are logged values and dark traces are centered nine-point moving averages. Dashed lines mark the evaluated step-600 checkpoints.
Figure 6: MoCo Rubric vs. Qwen3-VL-8B-Instruct. The selected examples illustrate differences in object pose, leaf shape, product layout, text alignment, and label position.
Figure 7: MoCo Rubric vs. RubiCap. The selected examples illustrate fabricated facial blurring and errors in text layout and typography.
Figure 8: MoCo Rubric vs. EvoLM. The selected examples illustrate fabricated facial blurring and border details, as well as errors in text layout and transcription.
Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large language models (MLLMs). In pursuit of ever more detailed and accurate captions, recent work has increasingly turned to reinforcement learning (RL). However, existing captioning-RL methods and evaluation metrics often emphasize a narrow notion of caption quality, inducing trade-offs across core dimensions of captioning. For example, utility-oriented objectives can encourage noisy, hallucinated, or overlong captions that improve downstream question answering while harming fluency, whereas arena-style objectives can favor fluent but generic descriptions with limited usefulness. To address this, we propose a more balanced RL framework that jointly optimizes utility-aware correctness, reference coverage, and linguistic quality. In order to effectively optimize the resulting continuous multi-objective reward formulation, we apply GDPO-style reward-decoupled normalization to continuous-valued captioning rewards and show that it improves performance over vanilla GRPO. Additionally, we introduce length-conditional reward masking, yielding a more suitable length penalty for captioning. Across LLaVA-1.5-7B and Qwen2.5-VL 3B and 7B base models, our method consistently improves caption quality, with peak gains of +13.6 DCScore, +9.0 CaptionQA, and +29.0 CapArena across different models.
In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs generally fall into two categories: holistic response-level judgment across heterogeneous criteria, or alignment-based evaluation against reference captions. However, both paradigms suffer from fundamental limitations. Holistic rewards struggle to ensure factual accuracy and are prone to stylistic reward hacking, while reference-based rewards overly rely on rigid textual alignment, failing to preserve the completeness and diversity inherent to open-ended generation tasks. To address these challenges, CuRe reformulates reward modeling as fine-grained claim-level verification. Specifically, CuRe decomposes captions into category-aware atomic claims through a structured rubric, converting holistic evaluation into simpler and more reliable claim-level verification.
Mingqi Gao, Hongyuan Dong, Yifei Chen +6
1Tsinghua Shenzhen International Graduate School, Tsinghua University · 3LLM Department, Tencent · University of Chinese Academy of Sciences
Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key challenge, however, lies in designing effective and interpretable rewards. Existing methods often rely on either composite metrics (e.g., CLIP, OCR, and realism scores) with fixed weights or a single scalar reward distilled from human preference models, which can limit interpretability and flexibility. We propose RubricRL, a simple and general framework for rubric-based reward design that offers greater interpretability, composability, and user control. Instead of using a black-box scalar signal, RubricRL dynamically constructs a structured rubric for each prompt--a decomposable checklist of fine-grained visual criteria such as object correctness, attribute accuracy, OCR fidelity, and realism--tailored to the input text. Each criterion is independently evaluated by a multimodal judge (e.g., o4-mini), and a prompt-adaptive weighting mechanism emphasizes the most relevant dimensions. This design not only produces interpretable and modular supervision signals for policy optimization (e.g., GRPO or PPO), but also enables users to directly adjust which aspects to reward or penalize. Experiments with an autoregressive text-to-image model demonstrate that RubricRL improves prompt faithfulness, visual detail, and generalizability, while offering a flexible and extensible foundation for interpretable RL alignment across text-to-image architectures.
Xuelu Feng, Yunsheng Li, Ziyu Wan +4
University at Buffalo · Microsoft AI · Nikola Tesla STEM High School