Organizations: State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Zhongguancun Academy
Object hallucination remains a major challenge for large vision-language models. While off-policy preference optimization proves to be an effective solution, on-policy reinforcement learning provides a more promising direction as it directly targets a model's current failure modes. However, we find that without fine-grained reward formulation and allocation, on-policy optimization often falls into an easy shortcut: reducing hallucinations merely by saying less---making fewer valid claims. To comprehensively resolve this, we propose a fine-grained alignment framework that couples dense reward signals at the data level with precise credit assignment at the algorithmic level. Specifically, we first construct the Dense Object Presence and Absence (DOPA) dataset to address sparse annotations that prevent valid object claims from being verified and rewarded. DOPA exhaustively annotates the deterministic presence and absence of every concept across an expanded vocabulary, significantly increasing the density of reliable reward signals during on-policy rollouts. Second, we propose Subsentence-level Credit Assignment for on-Policy Optimization (SCAPO) to prevent response-level shared advantages from allowing local hallucinations to compromise all other valid outputs within the same response. By assigning credit to each subsentence independently based on its object claims, SCAPO can precisely reinforce faithful generations and penalize hallucinations. Furthermore, we leverage the resulting faithful image descriptions as auxiliary context to transfer generative gains to discriminative tasks. Experiments demonstrate that our method produces highly informative, faithful descriptions in generative tasks while yielding clear performance gains on discriminative evaluation.
Figures & tables
Figure 1: Motivation for our approach. Conventional optimization suffers from sparse annotations leaving valid claims unrewarded, and response-level advantages allowing local hallucinations to suppress correct outputs. By combining the dense rewards of DOPA with the subsentence-level credit assignment of SCAPO, we achieve fine-grained alignment, preventing this suppression and yielding informative and faithful descriptions.
Figure 2: Pipeline for constructing the DOPA dataset. The process begins with rollout-driven vocabulary construction to establish the canonical vocabulary V and alias mapping A . Based on these, large-scale coarse annotation and edge-case refinement are performed. Ultimately, a rigorous manual verification cycle yields the high-fidelity object-existence matrix E∗ .
Figure 3: Overview of the SCAPO method and the adopted inference paradigm. Current-policy rollouts are segmented at subsentence boundaries and evaluated against the DOPA reward oracle. The resulting subsentence-level rewards are independently assigned to their corresponding token spans for the on-policy update. Furthermore, the optimized model leverages its generated faithful descriptions as auxiliary context for description-augmented discriminative inference.
AMBER
MS COCO
PhD
Method
Cap. Score ↑
Dis. Acc. ↑
Cap. Score ↑
Cap. Score † ↑
PhD-base ↑
PhD-sec ↑
PhD-icc ↑
PhD-ccs ↑
PhD-all ↑
(i) Inference-time Methods
Qwen2.5-VL
76.5
83.9
73.2
64.2
76.8
68.2
62.1
59.5
66.6
VCD
76.5
83.8
72.7
65.9
77.0
71.1
63.6
65.2
69.2
OPERA
76.3
84.5
73.7
65.3
76.4
70.9
64.6
62.0
68.5
VTI
47.7
85.9
50.2
33.6
79.0
57.3
58.3
66.8
65.4
Table 1: Evaluation on generative and discriminative hallucination benchmarks. Cap. Score = HarMean( 1-Hal. Rate, Cover Rate ) jointly measures caption faithfulness and informativeness. (Cap. Score is computed with the original object-existence annotation of the dataset. Cap. Score † is computed with our expanded DOPA annotation.)
MS COCO Label
DOPA Label
AMBER
MMHal-Bench
Method
Hal. Rate ↓
Cover Rate ↑
Hal. Rate ↓
Cover Rate ↑
Hal. Rate ↓
Cover Rate ↑
Score ↑
Hal. Rate ↓
Qwen2.5-VL (Baseline)
15.1
64.3
11.8
50.5
4.9
64.0
3.26
39.6
SCAPO Training
w/ MS COCO annotations
13.9
78.8
11.6
52.0
4.6
66.1
3.45
34.4
w/ DOPA annotations ( Ours )
13.4
73.7
9.4
69.6
4.8
73.2
3.52
33.3
Table 2: Effect of the DOPA dataset used for on-policy training. Both variants below utilize the same SCAPO objective but derive rewards from different annotations.
Annotated Objects per Response
Subsentences per Response
Annotation
Avg.
Sup.
Hal.
Avg.
Sup.
Hal.
MS COCO
2.73
2.32
0.41
6.33
2.62
0.45
DOPA ( Ours )
5.13
4.37
0.76
6.37
3.46
0.81
Table 3: Annotation density on training-set rollouts. For each annotation setting, we sample eight responses per training example and report the per-response averages. “Sup.” and “Hal.” denote ground-truth-supported and hallucinated content, respectively.
Figure 4: Training dynamics under different credit-assignment strategies. Response-level objectives face a trade-off between limited coverage (when optimizing Hal. Rate) and increased hallucinations (when optimizing Cap. Score). Object-level assignment fragments semantic coherence, causing unstable increases in both metrics. In contrast, SCAPO’s subsentence-level credit assignment resolves these issues, steadily expanding valid coverage while keeping the hallucination rate controlled.
Figure 5: FaithScore and Avg. facts across four generation settings. GRPO-style methods all improve faithfulness at the expense of the number of atomic facts in the output, while SCAPO preserves factual richness and maintains high faithfulness.
Configuration
Acc. ↑
Pre. ↑
Rec. ↑
F1 ↑
Qwen2.5-VL
83.9
84.5
92.6
88.4
+ Desc. Aug.
86.9 (+3.0)
87.3 (+2.8)
93.8 (+1.2)
90.4 (+2.0)
+ SCAPO
87.4 (+3.5)
87.6 (+3.1)
94.4 (+1.8)
90.9 (+2.5)
Table 4: Progressive ablation on the AMBER discriminative task. We first augment the Qwen2.5-VL baseline with a self-generated image description, and then upgrade the underlying backbone to the SCAPO-trained policy to evaluate the impact of generation quality.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Value
Number of images
5,000
Canonical object labels ( ∣V∣ )
160
Aliases ( ∣A∣ )
894
Average aliases per canonical label
5.59
Average positive annotations per image
10.75
Appendix
Table 5: Overall statistics of the DOPA dataset.
Number of positive images
# Labels
Percentage
<100
76
47.5%
100 – 499
59
36.9%
500 – 999
11
6.9%
≥1,000
14
8.8%
Total
160
100%
Appendix
Table 6: Label distribution according to the number of images in which they are annotated as present.
1
πθold←πθ // freeze the rollout policy
2
D←∅ // rollout buffer for this step
3
for each (x,p)∈B do
4
for i=1 to G do
5
yi∼πθold(⋅∣x,p)
6
[si,1,…,si,Mi]←\textscSplitSubsentences(yi)
7
{Ti,j}←\textscMapToTokenSpans(yi,{si,j}) // spans partition the response
Appendix
Figure 6: Pseudocode for one SCAPO optimization step.
Setting
Value
Model
Backbone
Qwen2.5-VL-7B-Instruct
Tuning scheme
Full fine-tuning
Vision tower
Frozen
Parameter dtype
bfloat16
Training
Appendix
Table 7: SCAPO training configuration. Values apply to all reported SCAPO results unless stated otherwise.
MME
MMBench
MMMU
LLaVA-Bench
MM-Vet
Method
Percep. ↑
Cognit. ↑
Circular ↑
Vanilla ↑
Acc. ↑
Rel. Score ↑
Total ↑
Qwen2.5-VL
1717.8
625.7
84.8
88.0
53.8
69.7
65.9
+ SCAPO (Ours)
1728.2
635.4
84.4
88.2
53.8
70.4
66.0
Appendix
Table 8: Results on general multimodal benchmarks.
Variant
Avg. Score ↑
Hal. Rate ↓
Qwen2.5-VL
3.26
39.6
Additive hallucination penalty
3.48
33.3
Full reward for repeated claims
3.41
37.5
Without no-object regularization
3.46
34.4
SCAPO (Ours)
3.52
33.3
Appendix
Table 9: Ablation of SCAPO’s components.
AMBER
MS COCO
PhD
Method
Cap. Score ↑
Dis. Acc. ↑
Cap. Score ↑
Cap. Score † ↑
PhD-all ↑
Qwen3-VL
76.0
86.6
72.8
67.1
75.0
+ SCAPO (Ours)
78.8
87.6
76.2
73.1
75.8
Appendix
Table 10: Results with Qwen3-VL backbone. Cap. Score † is computed against the expanded DOPA annotations.
Figure 7: Qualitative comparison of description-augmented discriminative inference between the baseline (Qwen2.5-VL) and our SCAPO-aligned model.
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
Qilu University of Technology (Shandong Academy of Sciences) Jinan, China · China Telecom Digital Intelligence Technology Co, Ltd Jinan, China · Shenyang Aerospace University Shenyang, China +1
School of Astronautics, Beihang University, Beijing 102206, China · Longcat Interaction Team, Meituan, Beijing 100102, China · Tianmushan Laboratory, Beihang University, Hangzhou 311115, China