Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control of false positives at the image level. To address this, we propose False Discovery Rate-COntRol of HALlucination (CORAL), a training-free framework that models visual uncertainty using an uncertainty-aware visual data splitting strategy and leverages mirror statistics to quantify visual contrast during decoding. By computing mirror statistics from paired, symmetrically perturbed visual inputs, CORAL estimates spurious object predictions and sets a data-driven threshold to control the expected fraction of false discoveries per image, suppressing hallucinations while retaining high power for truly grounded objects. The framework is flexible, supports multiple LVLMs, and mitigates hallucinations without retraining or supervision. Extensive experiments on multiple benchmarks with several evaluation metrics demonstrate that CORAL consistently outperforms state-of-the-art methods, providing more reliable and robust hallucination control. Code is available at: https://changliu1993-cl.github.io/CORAL/
Figures & tables
Figure 1 : Hallucination under multiple queries. The model may generate plausible but hallucinated objects (e.g., “shrimp” , in red ). CORAL performs image-level selection, rejecting hallucinations while retaining visually grounded objects under False Discovery Rate (FDR) control.
Figure 2 : Illustration of CORAL , which applies visual uncertainty splitting and mirror statistic to control the FDR of hallucinated objects. The hallucinated object “Dog” ( red ) is removed under FDR control.
Figure 3 : Uncertainty-aware visual data splitting . Paired, symmetrically perturbed visual views are generated from a shared noise source, producing contrastive views for the mirror statistic construction and thereby facilitating the separation of visual signal from noise.
Method
LLaVA-OneVision-7B
Qwen2.5-VL-7B
InternVL3-8B
Average
Overall FDR ↓
Overall Power ↑
Overall FDR ↓
Overall Power ↑
Overall FDR ↓
Overall Power ↑
Overall FDR ↓
Overall Power ↑
Random
Regular
0.0935 ( ± 0.0133)
75.36 ( ± 1.88)
0.0944 ( ± 0.0114)
77.43 ( ± 1.22)
0.0946 ( ± 0.0122)
73.52 ( ± 1.77)
0.0942 ( ± 0.0111)
75.44 ( ± 1.44)
VCD ( 2024 )
0.0892 ( ± 0.0188)
78.50 ( ± 1.23)
0.0908 ( ± 0.0132)
79.14 ( ± 1.53)
0.0899 ( ± 0.0109)
78.86 ( ± 1.34)
0.0900 ( ± 0.0145)
78.83 ( ± 1.31)
MARINE ( 2025 )
0.0846 ( ± 0.0122)
81.25 ( ± 1.17)
0.0866 ( ± 0.0116)
80.05 ( ± 1.60)
0.0889 ( ± 0.0133)
79.84 ( ± 1.33)
0.0867 ( ± 0.0124)
80.38 ( ± 1.21)
AGLA ( 2025 )
0.0799 ( ± 0.0223)
84.92 ( ± 1.11)
0.0821 ( ± 0.0114)
84.67 ( ± 1.41)
0.0810 ( ± 0.0144)
80.92 ( ± 1.33)
0.0810 ( ± 0.0123)
83.50 ( ± 1.50)
Table 1 : Evaluation of overall FDR control and overall power across multiple LVLM architectures on MSCOCO (3000 runs). Overall FDR and power denote averages of per-image FDR and power. Bold indicates the best result and underline indicates the second-best result.
Method
LLaVA-OneVision-7B
Qwen2.5-VL-7B
InternVL3-8B
Average
Accuracy ↑
F1 Score ↑
Accuracy ↑
F1 Score ↑
Accuracy ↑
F1 Score ↑
Accuracy ↑
F1 Score ↑
Random
Regular
85.87 ( ± 1.77)
85.72 ( ± 1.66)
86.20 ( ± 1.30)
86.98 ( ± 1.28)
87.15 ( ± 1.88)
86.26 ( ± 1.80)
86.41 ( ± 1.65)
86.32 ( ± 1.58)
VCD ( 2024 )
86.33 ( ± 1.23)
88.86 ( ± 1.43)
87.35 ( ± 1.83)
90.06 ( ± 2.03)
88.15 ( ± 1.53)
88.05 ( ± 1.45)
87.28 ( ± 1.53)
88.99 ( ± 1.64)
MARINE ( 2025 )
87.21 ( ± 1.35)
89.09 ( ± 1.09)
89.72 ( ± 1.11)
90.33 ( ± 2.09)
89.23 ( ± 1.35)
89.15 ( ± 1.19)
88.72 ( ± 1.27)
89.52 ( ± 1.46)
AGLA ( 2025 )
88.15 ( ± 1.65)
89.97 ( ± 1.21)
90.01 ( ± 1.23)
91.17 ( ± 1.51)
92.53 ( ± 1.42)
93.54 ( ± 1.77)
90.23 ( ± 1.43)
91.56 ( ± 1.50)
Table 2 : Evaluation with POPE score across multiple LVLM architectures on the MSCOCO dataset. We report individualized Accuracy and F1 score (mean ± std over 3000 runs). Bold indicates the best result and underline indicates the second-best result.
Model
Method
Object
Attribute
Relation
Total ↑
Existence ↑
Count ↑
Color ↑
Position ↑
Commonsense ↑
LLaVA-OneVision-7B
Regular
190.33 ( ± 6.50)
145.53 ( ± 15.20)
170.66 ( ± 9.10)
160.25 ( ± 8.30)
70.16 ( ± 6.50)
736.93 ( ± 24.80)
VCD ( 2024 )
186.25 ( ± 7.22)
147.25 ( ± 11.44)
175.35 ( ± 15.58)
165.36 ( ± 2.55)
79.42 ( ± 5.74)
753.63 ( ± 18.76)
MARINE ( 2025 )
191.26 ( ± 4.55)
150.44 ( ± 10.15)
177.45 ( ± 10.96)
165.36 ( ± 7.19)
82.33 ( ± 5.29)
766.84 ( ± 17.47)
AGLA ( 2025 )
190.54 ( ± 6.22)
153.32 ( ± 15.37)
178.42 ( ± 3.22)
170.35 ( ± 4.22)
84.36 ( ± 5.21)
776.99 ( ± 15.96)
CORAL (Ours)
193.67 ( ± 3.21)
160.43 ( ± 10.11)
180.35 ( ± 8.46)
175.24 ( ± 10.58)
85.52 ( ± 4.22)
795.21 ( ± 18.12)
Table 3 : Evaluation with MME score (mean ± std over 3000 runs) across multiple LVLM architectures on MSCOCO. Bold indicates the best result and underline indicates the second-best result.
Method
LLaVA-OneVision-7B
Qwen2.5-VL-7B
InternVL3-8B
Regular
80.77 ( ± 1.21)
83.50 ( ± 1.01)
83.40 ( ± 0.85)
VCD ( 2024 )
82.67 ( ± 1.00)
85.14 ( ± 1.32)
86.84 ( ± 0.99)
MARINE ( 2025 )
85.31 ( ± 1.09)
87.44 ( ± 1.17)
88.67 ( ± 1.03)
AGLA ( 2025 )
85.76 ( ± 1.11)
87.02 ( ± 1.07)
90.15 ( ± 1.00)
CORAL (Ours)
86.70 ( ± 1.13)
87.90 ( ± 0.91)
94.50 ( ± 0.80)
Table 4 : Evaluation with MMBench score (mean ± std over 3000 runs) across multiple LVLM architectures on MSCOCO. Bold indicates the best result and underline indicates the second-best result.
Method
Accuracy ↑
Precision ↑
Recall ↑
F1 Score ↑
Regular
85.87 ( ± 1.77)
83.41 ( ± 2.13)
88.08 ( ± 1.47)
85.72 ( ± 1.66)
w/o Visual Uncertainty Splitting
88.12 ( ± 0.62)
86.78 ( ± 0.74)
89.07 ( ± 0.71)
87.90 ( ± 0.58)
w/o Mirror Statistic
88.76 ( ± 0.55)
87.07 ( ± 0.49)
90.33 ( ± 0.61)
88.67 ( ± 0.53)
w/o Overall FDR Control
89.36 ( ± 0.48)
86.42 ( ± 0.92)
92.37 ( ± 0.56)
89.27 ( ± 0.41)
CORAL (Ours)
92.17 ( ± 0.89)
88.25 ( ± 1.46)
94.87 ( ± 1.61)
91.64 ( ± 1.83)
Table 5 : Ablation study on POPE metrics using the MSCOCO dataset with LLaVA-OneVision-7B. Results are mean ± std over 10 runs. Bold indicates the best result.
Figure 4 : Ablation study on the effect of FDR target level ( q ) on the performance of LLaVA-OneVision-7B, Qwen2.5-VL-7B, InternVL3-8B using POPE metrics with q={0.01,0.03,0.05,0.1,0.2} .
Figure 5 : Hallucination mitigation examples by our method CORAL across multiple tasks. Hallucinated objects are highlighted in red .
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Vision encoder
LLM
LLaVA-v1.5 [ 1 ]
CLIP-L-336px [ 57 ]
Vicuna-v1.5-7B [ 58 ]
Qwen-VL [ 53 ]
ViT-based visual encoder
Qwen-7B [ 53 ]
InstructBLIP [ 52 ]
BLIP-2 [ 3 ]
Vicuna-v1.1-7B [ 58 ]
LLaVA-OneVision-7B [ 49 ]
CLIP ViT-L/14 (336px) [ 57 ]
Vicuna-7B [ 58 ]
Qwen2.5-VL-7B [ 50 ]
ViT-based visual encoder
Qwen2.5-7B [ 50 ]
InternVL3-8B [ 51 ]
InternViT (high-resolution ViT)
InternLM2-8B [ 51 ]
Appendix
Table 6 : Details of the LVLM architectures used in our experiments.
Template Type
Prompt Template
POPE task
This image contains only the following objects: <OBJECT_GROUNDING> . Do not assume any objects beyond this list. Based solely on this information, <QUERY> The detected objects in the image are: <OBJECT_GROUNDING> . Answer the question using only these objects. <QUERY> This image shows the following objects: <OBJECT_GROUNDING> . You must answer using only the objects in this list.Given these detected objects, <QUERY> The objects found in this image are limited to: <OBJECT_GROUNDING> . You should rely strictly on this list of objects and make no other guesses. Based on this, <QUERY>
CORAL grounded
This image contains the following visually grounded objects: <OBJECT_GROUNDING> . Based on the image, <QUERY> The following objects are visible in the image: <OBJECT_GROUNDING> . Using only the information from the image, <QUERY> This image shows: <OBJECT_GROUNDING> . Please answer the following question based on the image content: <QUERY>
CORAL -restricted
The objects visible in this image are limited to: <OBJECT_GROUNDING> . Do not assume any objects beyond this list. <QUERY> Only the following objects appear in the image: <OBJECT_GROUNDING> . Answer the question using only visual evidence from the image. <QUERY> Based strictly on the objects shown in the image: <OBJECT_GROUNDING> . Do not infer any additional objects. <QUERY>
CORAL -complementary
The same image is analyzed using multiple internally constructed visual representations, including the original view and two mirrored views. Original view detects the following objects: <OBJECT_GROUNDING> Mirror view (+) detects the following objects: <OBJECT_GROUNDING_A> Mirror view (-) detects the following objects: <OBJECT_GROUNDING_B> Using the visual information above from the same image, <QUERY> Multiple complementary visual representations are derived from the same image. Original view objects: <OBJECT_GROUNDING> Mirrored view (+) objects: <OBJECT_GROUNDING_A> Mirrored view (-) objects: <OBJECT_GROUNDING_B> Based on the image, <QUERY>
Appendix
Table 7 : Details of the LVLM architectures that we used in our paper.
Figure 6 : Empirical validation of mirror symmetry for mirror statistic Δt . The figure shows overlaid histograms of Δt and its sign-flipped counterpart −Δt computed from negative POPE queries, where the queried object is guaranteed to be absent. The strong overlap between the two distributions and their near symmetry around zero indicate that mirror statistic for visually ungrounded objects are approximately symmetric, supporting their use as reference quantities for FDR estimation.
Figure 7 : Quantile-Quantile (QQ) plot comparing the empirical quantiles of Δt and −Δt for negative POPE queries. The close alignment with the identity line across the full range of quantiles indicates approximate distributional symmetry of the mirror statistic, providing empirical support for the mirror-symmetry assumption used in Eq. ( 8 ).
Parameters
Value
Amplification Factor α
1
Adaptive Plausibility Threshold β
0.1
Diffusion Noise Step
500
Appendix
Table 8 : VCD [ 24 ] Hyperparameter Settings.
Parameters
Value
Guidance Strength
0.7
Score Threshold for DERT
0.95
Detect Threshold for RAM++
0.68
Appendix
Table 9 : MARINE [ 22 ] Hyperparameter Settings.
Parameters
Value
Weighting Factor α
2
Adaptive Plausibility Constraint Factor β
0.5
Appendix
Table 10 : AGLA [ 23 ] Hyperparameter Settings.
Parameters
Value
Data Splitting Factor τ
0.1
FDR Thresholding q (false positive level in an image)
0.1
Appendix
Table 11 : CORAL Hyperparameter Settings. The settings are fixed depending on the question-answer tasks.
Model
Batch Size
LLaVA-v1.5
4
Qwen-VL
16
InstructBLIP
16
LLaVA-OneVision-7B
4
Qwen2.5-VL-7B
16
InternVL3-8B
8
Appendix
Table 12 : Batch size settings for LVLM generation across different models. Unless otherwise noted, the batch size is fixed for each model throughout all experiments. To improve evaluation efficiency, we adopt batched generation. When the LVLM does not explicitly specify a padding strategy for inference, we apply left padding to avoid potential negative effects of batched generation.
Method
LLaVA-v1.5
Qwen-VL
InstructBLIP
LLaVA-OneVision-7B
Qwen2.5-VL-7B
InternVL3-8B
CS↓
CI↓
Power ↑
CS↓
CI↓
Power ↑
CS↓
CI↓
Power ↑
CS↓
CI↓
Power ↑
CS↓
CI↓
Power ↑
CS↓
CI↓
Power ↑
Regular
9.2
5.1
94.8
9.5
19.2
80.7
8.7
4.9
95.2
8.8
4.6
95.4
8.2
18.1
81.9
5.0
3.2
96.8
VCD ( 2024 )
7.8
4.5
95.3
7.4
18.5
81.3
7.3
4.1
95.9
7.3
4.1
95.9
6.8
17.4
82.6
2.4
1.5
98.5
MARINE ( 2025 )
6.9
3.8
96.1
6.3
14.5
84.7
6.2
3.0
97.0
6.2
3.0
97.0
5.9
13.8
86.2
2.2
1.3
98.7
AGLA ( 2025 )
7.5
4.2
95.8
6.1
12.4
87.2
7.0
3.8
96.2
7.0
3.8
96.2
5.6
11.2
88.8
2.3
1.6
98.4
CORAL (Ours)
5.8
3.1
96.9
4.5
11.0
88.9
5.0
2.6
97.4
5.0
2.6
97.4
3.8
10.8
89.2
1.8
1.3
98.7
Appendix
Table 13 : Evaluation with CHAIR score across multiple LVLM architectures. Lower CS and CI indicate fewer hallucinated objects, while higher power indicates better retention of visually grounded objects. Bold indicates the best result.
Model
Method
Attribute
Relation
Total ↑
Existence ↑
Count ↑
Color ↑
Position ↑
LLaVA-v1.5
Regular
175.67 ( ± 7.51)
124.67 ( ± 19.59)
151.00 ( ± 10.45)
114.00 ( ± 9.32)
565.33 ( ± 33.92)
VCD ( 2024 )
184.66 ( ± 6.81)
138.33 ( ± 15.68)
153.00 ( ± 7.58)
128.67 ( ± 7.21)
604.66 ( ± 18.76)
MARINE ( 2025 )
190.53 ( ± 7.26)
154.43 ( ± 16.01)
166.34 ( ± 6.96)
130.28 ( ± 8.01)
641.58 ( ± 17.47)
AGLA ( 2025 )
195.00 ( ± 7.32)
153.89 ( ± 16.32)
167.67 ( ± 6.42)
129.44 ( ± 7.81)
646.00 ( ± 15.96)
CORAL (Ours)
194.43 ( ± 8.38)
157.41 ( ± 15.11)
167.34 ( ± 7.33)
131.19 ( ± 7.58)
650.37 ( ± 18.12)
Appendix
Table 14 : Evaluation with MME score across multiple LVLM architectures on MSCOCO with 3000 random replications. We report the mean ± standard deviation. Bold indicates the best result and underline indicates the second-best result.
Figure 8 : Inference latency comparisons. All measurements are conducted on a single NVIDIA RTX 6000 GPU with batch size 1, using identical input images and prompts.
Method
LLaVA-v1.5
Qwen-VL
InstructBLIP
Average
Overall FDR ↓
Overall Power ↑
Overall FDR ↓
Overall Power ↑
Overall FDR ↓
Overall Power ↑
Overall FDR ↓
Overall Power ↑
Random
Regular
0.0985 ( ± 0.0144)
77.32 ( ± 1.41)
0.0913 ( ± 0.0122)
76.08 ( ± 1.23)
0.0957 ( ± 0.0102)
74.35 ( ± 1.63)
0.0952 ( ± 0.0122)
75.92 ( ± 1.42)
VCD ( 2024 )
0.0965 ( ± 0.0200)
82.50 ( ± 1.13)
0.0898 ( ± 0.0111)
78.10 ( ± 1.12)
0.0941 ( ± 0.0119)
76.80 ( ± 1.22)
0.0935 ( ± 0.0143)
79.13 ( ± 1.16)
MARINE ( 2025 )
0.0935 ( ± 0.0112)
89.20 ( ± 1.07)
0.0852 ( ± 0.0119)
79.05 ( ± 1.13)
0.0944 ( ± 0.0114)
77.14 ( ± 1.35)
0.0910 ( ± 0.0115)
81.80 ( ± 1.18)
AGLA ( 2025 )
0.0918 ( ± 0.0107)
92.80 ( ± 1.15)
0.0818 ( ± 0.0132)
82.67 ( ± 1.10)
0.0921 ( ± 0.0122)
78.92 ( ± 1.21)
0.0886 ( ± 0.0120)
84.80 ( ± 1.15)
Appendix
Table 15 : Evaluation of overall FDR control and overall power across multiple LVLM architectures on MSCOCO with 3000 random replications. Overall FDR and power denote averages of per-image FDR and power across images. Bold indicates the best result and underline indicates the second-best result in each column.
Method
LLaVA-v1.5
Qwen-VL
InstructBLIP
Average
FDR ↓
Power ↑
FDR ↓
Power ↑
FDR ↓
Power ↑
FDR ↓
Power ↑
Random
Regular
0.0957 ( ± 0.0131)
77.57 ( ± 1.28)
0.0953 ( ± 0.0133)
79.53 ( ± 1.52)
0.0987 ( ± 0.0107)
75.34 ( ± 1.26)
0.0966 ( ± 0.0124)
77.48 ( ± 1.35)
VCD ( 2024 )
0.0930 ( ± 0.0100)
84.10 ( ± 1.22)
0.0892 ( ± 0.0121)
82.30 ( ± 1.13)
0.0967 ( ± 0.0112)
80.40 ( ± 1.15)
0.0930 ( ± 0.0111)
82.27 ( ± 1.17)
MARINE ( 2025 )
0.0920 ( ± 0.0101)
85.02 ( ± 1.35)
0.0835 ( ± 0.0202)
86.64 ( ± 1.05)
0.0920 ( ± 0.0100)
84.93 ( ± 1.08)
0.0892 ( ± 0.0134)
85.53 ( ± 1.16)
AGLA ( 2025 )
0.0892 ( ± 0.0112)
86.42 ( ± 1.42)
0.0832 ( ± 0.0158)
87.01 ( ± 1.11)
0.0923 ( ± 0.0123)
87.43 ( ± 1.11)
0.0882 ( ± 0.0131)
86.95 ( ± 1.21)
Appendix
Table 16 : Evaluation on overall FDR control and overall power across multiple LVLM architectures on A-OKVQA with 3000 random replications. Bold indicates the best result and underline indicates the second-best result in each column.
Method
LLaVA-v1.5
Qwen-VL
InstructBLIP
Average
FDR ↓
Power ↑
FDR ↓
Power ↑
FDR ↓
Power ↑
FDR ↓
Power ↑
Random
Regular
0.0933 ( ± 0.0132)
76.55 ( ± 1.32)
0.0914 ( ± 0.0188)
81.66 ( ± 1.01)
0.0889 ( ± 0.0188)
83.88 ( ± 1.52)
0.0912 ( ± 0.0169)
80.69 ( ± 1.28)
VCD ( 2024 )
0.0924 ( ± 0.0098)
80.92 ( ± 0.68)
0.0924 ( ± 0.0090)
80.24 ( ± 1.00)
0.0872 ( ± 0.0145)
85.47 ( ± 0.62)
0.0907 ( ± 0.0111)
82.21 ( ± 0.77)
MARINE ( 2025 )
0.0894 ( ± 0.0128)
86.92 ( ± 0.40)
0.0902 ( ± 0.0101)
84.22 ( ± 0.93)
0.0837 ( ± 0.0144)
87.39 ( ± 0.55)
0.0878 ( ± 0.0124)
86.18 ( ± 0.63)
AGLA ( 2025 )
0.0802 ( ± 0.0164)
90.42 ( ± 1.02)
0.0872 ( ± 0.0100)
85.11 ( ± 1.09)
0.0855 ( ± 0.0148)
89.43 ( ± 0.41)
0.0843 ( ± 0.0137)
88.99 ( ± 0.84)
Appendix
Table 17 : Evaluation on overall FDR control and overall power across multiple LVLM architectures on GQA with 3000 random replications. Bold indicates the best result and underline indicates the second-best result in each column.
Decoding
LLaVA-v1.5
Qwen-VL
Accuracy ↑
Precision ↑
Recall ↑
F1 Score ↑
Accuracy ↑
Precision ↑
Recall ↑
F1 Score ↑
Random
Regular
83.29 ( ± 0.35)
92.13 ( ± 0.54)
72.80 ( ± 0.57)
81.33 ( ± 0.41)
84.73 ( ± 0.36)
95.61 ( ± 0.45)
72.81 ( ± 0.38)
82.67 ( ± 0.41)
VCD ( 2024 )
87.73 ( ± 0.40)
91.42 ( ± 0.55)
83.28 ( ± 0.42)
87.16 ( ± 0.41)
88.63 ( ± 0.10)
94.64 ( ± 0.25)
81.91 ( ± 0.19)
87.81 ( ± 0.11)
MARINE ( 2025 )
85.01 ( ± 0.24)
88.27 ( ± 0.83)
80.73 ( ± 0.12)
84.33 ( ± 0.31)
82.07 ( ± 0.14)
89.27 ( ± 0.13)
89.33 ( ± 0.24)
85.83 ( ± 0.87)
AGLA ( 2025 )
88.54 ( ± 0.64)
94.41 ( ± 0.50)
82.08 ( ± 0.47)
87.71 ( ± 0.51)
84.60 ( ± 0.76)
98.23 ( ± 0.29)
70.47 ( ± 0.41)
82.07 ( ± 0.11)
Appendix
Table 18 : Evaluation with POPE score across multiple LVLM architectures on MSCOCO dataset comparing our method with several baselines. Bold indicates the best result and underline indicates the second-best result in each column.
Decoding
LLaVA-v1.5
Qwen-VL
Accuracy ↑
Precision ↑
Recall ↑
F1 Score ↑
Accuracy ↑
Precision ↑
Recall ↑
F1 Score ↑
Random
Regular
83.45 ( ± 0.48)
87.24 ( ± 0.68)
78.36 ( ± 0.54)
82.56 ( ± 0.50)
86.67 ( ± 0.48)
93.16 ( ± 0.55)
79.16 ( ± 0.59)
85.59 ( ± 0.53)
VCD ( 2024 )
86.15 ( ± 0.23)
85.18 ( ± 0.34)
87.53 ( ± 0.14)
86.34 ( ± 0.21)
89.22 ( ± 0.14)
90.77 ( ± 0.04)
87.32 ( ± 0.34)
89.01 ( ± 0.16)
MARINE ( 2025 )
86.72 ( ± 0.14)
87.71 ( ± 0.53)
87.53 ( ± 0.34)
86.34 ( ± 0.56)
89.17 ( ± 0.45)
89.87 ( ± 0.42)
88.13 ( ± 0.54)
88.93 ( ± 0.18)
AGLA ( 2025 )
89.28 ( ± 0.33)
93.18 ( ± 0.43)
84.76 ( ± 0.42)
88.77 ( ± 0.56)
86.77 ( ± 0.43)
95.02 ( ± 0.11)
77.60 ( ± 0.72)
85.43 ( ± 0.98)
Appendix
Table 19 : Evaluation with POPE score across multiple LVLM architectures on A-OKVQA dataset comparing our method with several baselines with 3000 random replications. Bold indicates the best result and underline indicates the second-best result in each column.
Decoding
LLaVA-v1.5
Qwen-VL
Accuracy ↑
Precision ↑
Recall ↑
F1 Score ↑
Accuracy ↑
Precision ↑
Recall ↑
F1 Score ↑
Random
Regular
83.73 ( ± 0.27)
87.16 ( ± 0.39)
79.12 ( ± 0.35)
82.95 ( ± 0.28)
80.97 ( ± 0.32)
88.07 ( ± 0.34)
71.64 ( ± 0.57)
79.01 ( ± 0.40)
VCD ( 2024 )
86.65 ( ± 0.45)
84.58 ( ± 0.59)
89.24 ( ± 0.34)
86.99 ( ± 0.41)
85.59 ( ± 0.38)
86.88 ( ± 0.44)
83.84 ( ± 0.36)
85.33 ( ± 0.38)
MARINE ( 2025 )
86.33 ( ± 0.14)
85.02 ( ± 0.63)
88.23 ( ± 0.15)
87.24 ( ± 0.11)
85.54 ( ± 0.52)
87.35 ( ± 0.25)
87.26 ( ± 0.22)
86.63 ( ± 0.22)
AGLA ( 2025 )
86.46 ( ± 0.34)
85.84 ( ± 0.25)
87.31 ( ± 0.15)
86.57 ( ± 0.15)
83.90 ( ± 0.43)
93.05 ( ± 0.09)
73.26 ( ± 0.82)
81.98 ( ± 0.13)
Appendix
Table 20 : Evaluation with POPE score across multiple LVLM architectures on GQA dataset, comparing our method with several baselines with 3000 random replications. Bold indicates the best result and underline indicates the second-best result in each column.
Figure 9 : Ablation study on the effect of FDR target level ( q ) on the performance of LLaVA-v1.5, Qwen-VL, InstructBLIP using POPE metrics with q={0.01,0.03,0.05,0.1,0.2} .
Figure 10 : Sensitivity analysis of FDR and power with respect to the perturbation scale τv . We report empirical FDR and power under different values of τv∈{0.01,0.05,0.10,0.20,1.00} while fixing the target FDR level at q=0.1 . The results show that CORAL achieves a favorable trade-off at intermediate perturbation scales and remains robust across a wide range of τv values.
Figure 11 : Sensitivity analysis of POPE accuracy and F1 with respect to the perturbation scale τv . Performance is evaluated under different values of τv∈{0.01,0.05,0.10,0.20,1.00} while fixing the target FDR level at q=0.1 . The results indicate that CORAL is not sensitive to precise tuning of τv and achieves stable performance across a broad range of perturbation scales.
Figure 12 : Representative failure cases of CORAL on LLaVA-v1.5. Each example compares responses with and without CORAL to the same image and query. The cases illustrate remaining errors in object recognition and activity description, including failure to recognize a present cat, an unsupported traffic-light prediction, an inconsistent response concerning a knife, and confusion involving a surfboard and parasailing. Red and green text highlight incorrect and correct response components, respectively.
Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient visual attention. However, we find that both real and hallucinated objects receive equally strong visual attention in the model's mid-to-late layers, suggesting that the key issue may not be how much the model attends, but what it attends to and why. To this end, we decode the visual features of high-attention regions using Logit Lens, and observe that regions corresponding to real objects can be correctly decoded to the target object tokens, whereas those for hallucinated objects cannot. Building on this, we identify two hallucination mechanisms: (i) visual uncertainty, triggered by semantically similar or confusable regions; masking these regions eliminates the hallucination. (ii) contextual prior, triggered by strong co-occurrence priors; even when the initially attended region is masked, the hallucination persists and attention drifts to other regions. Based on these findings, we propose a simple yet effective training-free Detect-Mitigate framework comprising a Logit-Lens Consistency Check to detect hallucination and targeted remedies: High-Attention Regions Masking (HARM) for visual uncertainty hallucination, and Visual Evidence Enhanced Decoding (VEED) for contextual prior hallucination. Our approach achieves state-of-the-art results on multiple hallucination benchmarks. Code will be available.
Zichuan Wang, Songlin Yang, Bo Peng +4
School of Artificial Intelligence, University of Chinese Academy of Sciences · 3Hong Kong University of Science and Technology · 2New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Science
Object Hallucination in large vision-language models (LVLMs), where models generate non-factual content about input images, remains a critical barrier to their reliability in real-world applications. Existing mitigation strategies can be categorized into training-based and training-free methods. Training-based methods often achieve strong performance but are costly, requiring extensive computational resources, large-scale data, and time-consuming fine-tuning. Training-free approaches are particularly appealing due to their efficiency. However, existing training-free methods either require multiple decoding rounds, which adds computational overhead, or modify internal states in a model-specific way that risks degrading pretrained knowledge. We propose Test-Time Hallucination Mitigation (TTH) method, a novel training-free method that addresses both limitations. TTH introduces a token-validator module, implemented as a zero-shot Multi-Modal Classifier (MMC), to generate auxiliary logits grounded in the input image. These logits are fused with the original LVLM outputs at the token level for object tokens selected from a candidate pool. An entropy-based weighting scheme is then applied to enable robust and accurate predictions. Extensive experiments across multiple LVLM families and diverse benchmarks demonstrate that TTH consistently improves accuracy and robustness, underscoring its generalizability and practical effectiveness. Code is released at https://github.com/Mehran-TAM/TTH
Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian +3
University of Technology Sydney · York University, Canada · Australian National University +2
The generation of factually incorrect objects, commonly known as object hallucination, remains a persistent challenge in Large Vision-Language Models (LVLMs). Current approaches to address this issue - ranging from expensive data-driven fine-tuning and high-latency contrastive decoding to rigid attention head truncation - frequently compromise either computational efficiency or the continuity of the model's feature space. To overcome these limitations, we introduce a novel, training-free inference strategy that operates as a region-aware adaptive weighting mechanism to dynamically correct semantic drift without relying on abrupt heuristic truncations. By computing an outlier-resistant statistical midpoint across various attention heads, we establish a stable anchor for reliable visual representations. We then utilize the inter-head disagreement mapped across regions to dynamically determine intervention budgets, gently suppressing hallucination-inducing attention paths through a continuous penalty modulation. This recalibration process effectively rectifies visual-semantic misalignments while fully preserving generative fluency and language priors. Comprehensive evaluations on standard multimodal benchmarks, including CHAIR, POPE, and MME, reveal that our strategy substantially curtails both instance- and sentence-level hallucinations. The results demonstrate state-of-the-art performance against contemporary baselines, confirming our method's efficiency and algorithmic robustness. Our code will be public.
Yuanzhi Xu, Qian Gao, Jun Fan +4
Qilu University of Technology (Shandong Academy of Sciences) Jinan, China · China Telecom Digital Intelligence Technology Co, Ltd Jinan, China · Shenyang Aerospace University Shenyang, China +1