Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision--language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.
Figures & tables
Figure 1: Comparison of heterogeneous and shared model distributions in rubric-based caption optimization.
Figure 2: Overview of MoCo Rubric. Stage I trains the three roles through shared multi-role SFT. In Stage II, current Policy candidates inform the rubric, the momentum Generator and Judge assign rewards, and only the Policy receives RL gradients. The shared rubric model is updated through EMA at synchronization boundaries.
Method
PixMo-Cap
DenseFusion
CapArena
CompreCap
DOCCI
Average
Captioning Baselines
ShareGPT4V-7B
2.02
1.80
2.67
5.21
2.41
2.82
RICO-Flash-7B
8.89
5.41
9.00
9.87
9.04
8.44
OmniCaptioner-8B
10.91
10.62
12.17
10.05
11.24
11.00
JoyCaption-8B
11.72
4.81
12.67
25.85
14.46
13.90
MetaCaptioner-8B
22.63
23.25
19.00
32.50
18.47
23.17
Table 1: Pairwise win rates (%) against Qwen3-VL-8B across five captioning benchmarks. Average denotes the benchmark mean. Best results are bold; second-best results are underlined.
Method
CaptionQA
BLINK
TextVQA
DocVQA
ChartQA
Average
Base
80.41
47.05
58.00
69.18
61.28
63.18
RubiCap
83.28
49.05
56.85
77.90
64.44
66.31
Qwen3-VL-32B
82.92
50.94
56.69
75.06
62.16
65.55
EvoLM
82.00
49.00
57.08
77.21
62.44
65.55
SFT (shared)
81.05
47.69
56.44
76.13
63.16
64.89
Ours
84.24
50.10
58.33
79.05
64.12
67.17
Table 2: Caption-based question answering across five benchmarks. Average is the mean across benchmarks. Bold and underlined values indicate the best and second-best results, respectively.
ID
SFT initialization
Rubrics
Rubric-role update
PixMo-Cap
DenseFusion
Average
M1
Policy only
—
—
57.37
56.91
57.14
M2
Shared across roles
—
—
55.15
53.31
54.23
M3
Separate per role
Offline, fixed
No momentum update
70.74
66.00
68.37
M4
Separate per role
Online
No momentum update
71.31
67.13
69.22
M5
Separate per role
Online
EMA ( m=0.99 , H=1 )
65.06
61.80
63.43
M6
Shared across roles
Online
No momentum update
68.48
64.73
66.61
Table 3: Component ablation on PixMo-Cap and DenseFusion, reported as win rates (%). Bold and underlined values indicate the best and second-best available results, respectively.
Figure 6
Figure 4: Left: Momentum coefficient analysis at H=1 , with tested coefficients equally spaced and the frozen reference ( m=1 ) shown separately. Middle: Synchronization interval analysis at m=0.99 . Markers indicate measured results; lines connect adjacent tested settings. Right: Policy–momentum tracking at H=1 . Frozen and full-sync curves are boundary references.
m
H=1
H=5
H=20
0.95
69.58
68.18
69.29
0.99
72.14
71.79
70.18
0.999
68.28
66.67
69.48
Table 5: Equal-weight mean win rate (%) across PixMo-Cap and DenseFusion in the joint m×H sweep.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
PixMo-Cap
DenseFusion
CapArena
CompreCap
DOCCI
Average
Captioning Baselines
ShareGPT4V-7B
1.00
0.20
0.50
1.25
1.00
0.79
RICO-Flash-7B
5.02
4.02
5.34
8.06
4.20
5.33
OmniCaptioner-8B
5.62
3.41
6.34
7.89
9.80
6.61
JoyCaption-8B
11.65
6.43
11.52
31.18
13.60
14.87
MetaCaptioner-8B
14.06
10.64
13.52
20.97
14.80
14.80
Appendix
Table 6: Pairwise win rates (%) against Qwen3-VL-8B with GPT-5.6 Sol as the evaluator. Average denotes the benchmark mean. Best results are bold; second-best results are underlined.
LLM judge
Exact agreement
Gemini 3.1 Pro
429/499 (85.97%)
GPT-5.6 Sol
412/497 (82.90%)
Appendix
Table 7: Agreement between LLM judges and final human preferences on Qwen3-8B versus GPT-5.6 caption pairs.
Figure 5: GRPO training dynamics of RubiCap and MoCo Rubric. The upper panels show mean training reward; the lower panels show the fraction of rollout groups with zero within-group reward standard deviation. Light traces are logged values and dark traces are centered nine-point moving averages. Dashed lines mark the evaluated step-600 checkpoints.
Figure 6: MoCo Rubric vs. Qwen3-VL-8B-Instruct. The selected examples illustrate differences in object pose, leaf shape, product layout, text alignment, and label position.
Figure 7: MoCo Rubric vs. RubiCap. The selected examples illustrate fabricated facial blurring and errors in text layout and typography.
Figure 8: MoCo Rubric vs. EvoLM. The selected examples illustrate fabricated facial blurring and border details, as well as errors in text layout and transcription.