Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At the rollout level, hard queries---those with high semantic entropy---frequently produce unanimously wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk is highest. At the optimization level, confident-but-wrong tokens are gradient-invisible: a categorical policy's expected score-gradient norm vanishes as its distribution sharpens, so the predictions that most need correction receive the weakest updates. We propose Dual-Entropy Enhanced Policy Optimization (DEEPO), a dual-stage enhancement combining signal variance regularization with gradient preconditioning: semantic-entropy-triggered expert prefixes inject grounded continuations on high-uncertainty queries, providing direct supervision and restoring advantage variance, while advantage-sign-aware Renyi preconditioning counteracts logit-level saturation so correction reaches confident errors in the operational confidence regime. Both branches improve over GRPO individually; their interaction is statistically significant on VideoMMMU---the most complex long-horizon task in our evaluation suite (+4.0$, 95% CI [1.1, 6.9])---and additive elsewhere. DEEPO reduces hallucination while preserving accuracy and training stability.
Figures & tables
Figure 1: Model guessing and expert-hint-guided correction. Under uncertain visual cues the model “guesses,” yielding high semantic entropy and frequently unanimous failures; the expert hint grounds one continuation on the visual evidence.
Figure 2: Overview of Dual-Entropy Enhanced Policy Optimization. High semantic entropy activates ϕt -weighted prefix regularization and prefix-guided rollouts, while the Rényi entropy factor scales policy gradients; the objective unifies both.
Model (Method)
POPE ↑
VideoHallucer ↑
HalluBench ↑
Closed-source Models
GPT-4o
86.9
53.3
55.0
Gemini-1.5-Pro
–
37.8
45.6
Open-source Multimodal Base Models
LLaVA-OneVision-7B
69.8 ± 1.5
44.6 ± 1.2
61.6 ± 1.1
InternVL2.5-8B
82.1 ± 0.6
50.5 ± 1.3
55.6 ± 0.7
Table 1: Results on six benchmarks ( ↑ : higher is better). Open-source rows are mean ± std over four runs (training seeds for fine-tuned models; bootstrap over examples for frozen baselines); closed-source rows are single-run. Bold: best among training strategies on the Qwen2.5-VL-7B backbone.
Model (Method)
POPE ↑
VideoHallucer ↑
HalluBench ↑
Qwen2.5-VL-7B
81.3
52.2
64.1
+ GRPO
83.7
52.0
68.6
+ w/o Expert Hint
84.4
54.2
68.7
+ w/o Gradient Scaling
84.7
55.2
68.2
+ DEEPO
85.8
57.6
69.7
Interaction I
+0.4
+0.2
+1.4
Table 2: Ablation of the two branches and their interaction ( ↑ : higher is better). All variants use Qwen2.5-VL-7B; all rows are four-seed means. “w/o Expert Hint” removes the whole signal-repair branch; “w/o Gradient Scaling” removes only the delivery branch, keeping the entropy-triggered prefixes, prefix regularizer, and dropout. The last row reports the interaction I=Sfull−Shint−Sscaling+SGRPO . The 95% CIs are [−0.4,1.2] / [−1.3,1.7] / [−0.3,3.1] / [−0.2,1.0] / [−0.8,1.2] / [1.1,6.9] (last bold); only VideoMMMU excludes zero.
Variant
POPE ↑
HallusionBench ↑
VideoMMMU ↑
GRPO
83.7
68.6
44.7
w/ symmetric- τ weights
85.0
69.2
50.4
w/ symmetric-inverse weights
85.2
69.3
51.1
w/ sign-only weights
84.9
69.0
50.0
w/ EMPG weights
85.3
69.4
51.5
w/ random trigger
85.1
69.3
51.4
Table 3: Delivery-branch design controls and trigger ablation (four-seed means, Qwen2.5-VL-7B); row 5 uses a rate-matched random trigger.
Hs decile
1
2
3
4
5
6
7
8
9
10
All-wrong (%)
12
16
21
27
33
39
45
51
57
63
Mean reward
0.76
0.69
0.61
0.53
0.46
0.40
0.34
0.29
0.24
0.19
Table 4: Mechanism diagnosis on the training distribution (four-seed pools): (a) all-wrong group fraction (%) and mean group reward by Hs decile (GRPO baseline); (b) hint activation on triggered, originally all-wrong queries; (c) baseline grouping ablation; (d) equal-compute and equal-supervision baselines. All rows in (c)–(d) run the identical rollout pipeline of Algorithm 1 ( G=8 unhinted rollouts plus one hint-conditioned continuation per triggered query); the “mixed” row is the default model itself, not a rerun, and in the “separated” variant the unhinted group is normalized within itself while the singleton hinted continuation is scored against a batch-level EMA baseline over hint continuations—so no advantage is zeroed by construction and only the counterfactual penalty channel is removed.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Effect of entropy-masked gradients. Masking low-entropy gradients preserves stable learning, while masking high-entropy gradients destabilizes training: high-entropy tokens carry the bulk of the update signal (Prop. 1 ); this pilot shows where gradient mass lives, not a hallucination outcome.
Figure 4: Semantic entropy on HallusionBench and MMStar (cloze-style conversion; questions sorted by the base model’s Hs for readability). Compared with the untrained base model, DAPO yields higher semantic entropy, while GRPO and PAPO reduce entropy but remain more variable. DEEPO maintains consistently lower and less dispersed semantic entropy, indicating reduced uncertainty-driven guessing.
Initial τs
POPE ↑
VideoHallucer ↑
HallusionBench ↑
MMBench ↑
MMSTAR ↑
VideoMMMU ↑
0.2
85.5
55.2
69.4
86.2
60.5
52.8
0.4
85.7
55.7
69.5
86.5
60.7
53.1
0.5
85.6
55.3
69.4
86.3
60.5
52.9
0.8
85.8
57.6
69.7
86.4
60.8
53.3
0.9
85.1
55.8
69.7
86.5
60.4
53.2
Appendix
Table 5: Sensitivity to the initial semantic entropy threshold τs (four-seed means).
Rollout N
POPE ↑
VideoHallucer ↑
HallusionBench ↑
MMBench ↑
MMSTAR ↑
VideoMMMU ↑
4
84.7
54.2
68.5
85.4
59.1
49.3
6
85.4
55.1
69.2
86.1
60.1
51.9
8
85.8
57.6
69.7
86.4
60.8
53.3
12
86.0
55.5
69.6
86.5
60.7
53.5
Appendix
Table 6: Sensitivity to the rollout budget N (co-varied with the GRPO group size G ; four-seed means).
Method
Hint
Trigger
POPE ↑
HallusionBench ↑
VideoMMMU ↑
GRPO
✗
✗
83.7
68.6
44.7
GRPO + Expert Hint
✓
✗
83.4
68.8
45.0
GRPO + Entropy-triggered Hint
✓
✓
84.2
69.0
46.7
DEEPO w/o Expert Hint
✗
✗
84.4
68.7
44.7
DEEPO
✓
✓
85.8
69.7
53.3
Appendix
Table 7: Controlled study on expert-hint supervision (four-seed means). Entropy-triggered hints improve over the corresponding base method, yet remain below DEEPO—hint supervision alone does not explain the gains. Hint rows use rollout conditioning only (no prefix regularizer, no dropout), unlike Table 2 ’s “w/o Gradient Scaling” row.
Variant
POPE ↑
HallusionBench ↑
VideoMMMU ↑
GRPO (four-seed mean)
83.7
68.6
44.7
DEEPO (shuffled hints)
83.2
67.7
44.5
DEEPO (correct hints, four-seed mean)
85.8
69.7
53.3
Appendix
Table 8: Hint-shuffle control (four-seed means): entropy-triggered prefixes from mismatched queries. Shuffled hints should fail to help despite identical context length and trigger statistics.
Variant
POPE ↑
HallusionBench ↑
VideoMMMU ↑
DEEPO (dataset hints, four-seed mean)
85.8
69.7
53.3
Deepo -self
84.6
68.6
50.8
Appendix
Table 9: Self-generated hints ( Deepo -self, four-seed means): prefixes come from the highest-reward verified-correct rollout within the same GRPO group; no dataset traces are used.
Benchmark
g⋆ (error)
g⋆ (non-error)
Enrichment ↑
95% CI
∥Δz∥ (err / non-err)
HallusionBench
1.68
0.96
1.75
[1.58,1.93]
0.31 / 0.27
VideoMMMU
1.80
1.00
1.80
[1.63,1.98]
0.35 / 0.28
Appendix
Table 10: Error-localization analysis on failed rollouts: per-token update intensity on hallucinated vs. non-error content. Annotation: token-level hallucinated-span labeling by two human annotators with adjudication (600 rollouts; Cohen’s κ=0.78 / 0.81 ), cross-checked by GPT-4V ( 92% token agreement on a 50% subsample). ∥Δz∥ : mean effective logit-step norm. CI: 95% bootstrap over rollouts.
Variant
POPE ↑
HallusionBench ↑
VideoMMMU ↑
InternVL2.5-8B
82.1
55.6
44.2
+ GRPO
83.4
57.0
45.3
+ DEEPO
85.0
58.2
51.0
Appendix
Table 11: Generalization to a second backbone (InternVL2.5-8B; four-seed means).
Model
HallusionBench ρ / ρsp
MMStar ρ / ρsp
GRPO (w/o hint branch)
0.48 / 0.45
0.44 / 0.42
DEEPO (w/ hint branch)
0.62 / 0.59
0.58 / 0.55
Appendix
Table 12: Correlation between Hˉ1 and Hs on multimodal benchmarks (Pearson ρ / Spearman ρsp ), for models trained without and with the hint branch. All coefficients significant at p<0.01 .
Figure 5: Effect of Gradient Scaling on RL stability. Adaptive weighting stabilizes RL loss, accelerates reward convergence, and reduces policy clipping, yielding smoother and more robust training.
Parameter
Value
Backbone
Qwen2.5-VL-7B Bai et al. (2025)
Training data
Video-R1 Feng et al. (2025) multimodal RL split, 20k subset
GRPO group size G
8
Rollout temperature
1.2
Optimizer (policy)
Adam, lr =5×10−7
Clip parameter εclip
0.2
Appendix
Table 13: Full hyperparameter configuration.
Method
Step Time (s) ↓
Wall-clock Δ
SE Time Share
MMBench ↑
Qwen2.5-VL-7B
-
-
-
83.1
+ GRPO
109.8 ± 2.7
1.00 ×
–
84.5
+ DEEPO w/o Expert Hint (scaling only)
111.9 ± 2.9
+1.9%
∼ 0.0%
85.0
+ DEEPO w/o Gradient Scaling (hint only)
123.6 ± 3.8
+12.6%
4.7% ± 0.6%
85.5
+ DEEPO full
129.6 ± 4.1
+18.0%
5.4% ± 0.7%
86.4
Appendix
Table 14: Efficiency–performance trade-off. “SE Time Share” covers answer extraction and clustering for Hs only; the conditional continuation pass and the prefix forward pass are separate delta components.
Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsistent with the visual input. Direct Preference Optimization (DPO) mitigates this by training models to prefer non-hallucinated responses over hallucinated ones, and recent efforts further enrich the preference data with relevant context. However, it remains unclear whether DPO actually leverages such context. To investigate this, we propose Contextual Preference Gain (CPG), a simple metric that measures how much a model's preference strengthens when relevant context is provided. We find that higher CPG consistently corresponds to lower hallucination, yet standard DPO and its variants exhibit only limited CPG, indicating that they underutilize contextual information and thus remain prone to hallucination. To address this, we propose Context-Calibrated DPO (C2-DPO), which directly maximizes CPG while preserving the original preference ordering. Across multiple benchmarks, C2-DPO substantially reduces hallucination without compromising general reasoning, relatively reducing the Object HalBench hallucination rate of Qwen2-VL-Instruct-2B by 36%. Code is available at https://github.com/mlvlab/C2-DPO
Byungoh Ko, Jinyoung Park, Jongha Kim +3
Korea University, Seoul, Republic of Korea · KAIST, Daejeon, Republic of Korea
Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) by learning from preference pairs. One of its key challenges lies in how to transfer the sequence-level preference into fine-grained supervision on visual fidelity. To safeguard vision-related tokens that are prone to hallucination, existing methods typically allocate training emphasis according to the model's self-assessed visual sensitivity signals. However, such sensitivity, estimated by a model still under training, introduces self-referential bias: reinforcing already well-learned visual cues while neglecting hard-to-perceive but critical details, thereby limiting deeper alignment. In this work, we propose an Uncertainty-aware Exploratory Direct Preference Optimization (UE-DPO) method for MLLMs, which enables the model to uncover its cognitive deficiencies and actively explore for self-correction, guided by token-level epistemic uncertainty. Specifically, we first quantify the uncertainty from the model's failure to ground token predictions in the given image. Then, based on an uncertainty-aware exploration intensity, we encourage more learning pressure on visually deficient tokens in preferred samples, and alleviate the over-penalization of beneficial knowledge in dispreferred samples. Further, we provide a theoretical justification for our method, and extensive experiments demonstrate its effectiveness and robustness.
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that are not grounded in the image. Many remedies either modify decoding at test time, which adds latency, or fine tune with preferences such as DPO variants, which teach which answer is preferred but not when the model's own answer is unreliable. We argue that calibrated self assessment is the missing signal. We introduce Savor, a training framework that (i) augments the output schema with token and answer confidence, (ii) optimises the policy with a Group Relative Policy Optimisation (GRPO) objective that penalises calibration error and poor abstention decisions, and (iii) uses the learned confidence at inference time to revisit visual evidence only when the model is uncertain. Experiments on POPE, HallusionBench, AMBER and MMHal-Bench across two recent backbones (InternVL3-8B and Qwen3-VL-8B) show that Savor reduces hallucination while preserving general capability on MME and MMBench, with lower Expected Calibration Error than DPO and decoding baselines.