Mitigating Object Hallucination in Large Vision-Language Models via False Discovery Controlled Visual Data Splitting
Organizations: School of Data, Mathematical, and Statistical Sciences Institute of Artificial Intelligence University of Central Florida
Abstract
Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control of false positives at the image level. To address this, we propose False Discovery Rate-COntRol of HALlucination (CORAL), a training-free framework that models visual uncertainty using an uncertainty-aware visual data splitting strategy and leverages mirror statistics to quantify visual contrast during decoding. By computing mirror statistics from paired, symmetrically perturbed visual inputs, CORAL estimates spurious object predictions and sets a data-driven threshold to control the expected fraction of false discoveries per image, suppressing hallucinations while retaining high power for truly grounded objects. The framework is flexible, supports multiple LVLMs, and mitigates hallucinations without retraining or supervision. Extensive experiments on multiple benchmarks with several evaluation metrics demonstrate that CORAL consistently outperforms state-of-the-art methods, providing more reliable and robust hallucination control. Code is available at: https://changliu1993-cl.github.io/CORAL/
Figures & tables
| Method | LLaVA-OneVision-7B | Qwen2.5-VL-7B | InternVL3-8B | Average | ||||
| Overall FDR | Overall Power | Overall FDR | Overall Power | Overall FDR | Overall Power | Overall FDR | Overall Power | |
| Random | ||||||||
| Regular | 0.0935 ( 0.0133) | 75.36 ( 1.88) | 0.0944 ( 0.0114) | 77.43 ( 1.22) | 0.0946 ( 0.0122) | 73.52 ( 1.77) | 0.0942 ( 0.0111) | 75.44 ( 1.44) |
| VCD ( 2024 ) | 0.0892 ( 0.0188) | 78.50 ( 1.23) | 0.0908 ( 0.0132) | 79.14 ( 1.53) | 0.0899 ( 0.0109) | 78.86 ( 1.34) | 0.0900 ( 0.0145) | 78.83 ( 1.31) |
| MARINE ( 2025 ) | 0.0846 ( 0.0122) | 81.25 ( 1.17) | 0.0866 ( 0.0116) | 80.05 ( 1.60) | 0.0889 ( 0.0133) | 79.84 ( 1.33) | 0.0867 ( 0.0124) | 80.38 ( 1.21) |
| AGLA ( 2025 ) | 0.0799 ( 0.0223) | 84.92 ( 1.11) | 0.0821 ( 0.0114) | 84.67 ( 1.41) | 0.0810 ( 0.0144) | 80.92 ( 1.33) | 0.0810 ( 0.0123) | 83.50 ( 1.50) |
| Method | LLaVA-OneVision-7B | Qwen2.5-VL-7B | InternVL3-8B | Average | ||||
| Accuracy | F1 Score | Accuracy | F1 Score | Accuracy | F1 Score | Accuracy | F1 Score | |
| Random | ||||||||
| Regular | 85.87 ( 1.77) | 85.72 ( 1.66) | 86.20 ( 1.30) | 86.98 ( 1.28) | 87.15 ( 1.88) | 86.26 ( 1.80) | 86.41 ( 1.65) | 86.32 ( 1.58) |
| VCD ( 2024 ) | 86.33 ( 1.23) | 88.86 ( 1.43) | 87.35 ( 1.83) | 90.06 ( 2.03) | 88.15 ( 1.53) | 88.05 ( 1.45) | 87.28 ( 1.53) | 88.99 ( 1.64) |
| MARINE ( 2025 ) | 87.21 ( 1.35) | 89.09 ( 1.09) | 89.72 ( 1.11) | 90.33 ( 2.09) | 89.23 ( 1.35) | 89.15 ( 1.19) | 88.72 ( 1.27) | 89.52 ( 1.46) |
| AGLA ( 2025 ) | 88.15 ( 1.65) | 89.97 ( 1.21) | 90.01 ( 1.23) | 91.17 ( 1.51) | 92.53 ( 1.42) | 93.54 ( 1.77) | 90.23 ( 1.43) | 91.56 ( 1.50) |
| Model | Method | Object | Attribute | Relation | Total | ||
| Existence | Count | Color | Position | Commonsense | |||
| LLaVA-OneVision-7B | Regular | 190.33 ( 6.50) | 145.53 ( 15.20) | 170.66 ( 9.10) | 160.25 ( 8.30) | 70.16 ( 6.50) | 736.93 ( 24.80) |
| VCD ( 2024 ) | 186.25 ( 7.22) | 147.25 ( 11.44) | 175.35 ( 15.58) | 165.36 ( 2.55) | 79.42 ( 5.74) | 753.63 ( 18.76) | |
| MARINE ( 2025 ) | 191.26 ( 4.55) | 150.44 ( 10.15) | 177.45 ( 10.96) | 165.36 ( 7.19) | 82.33 ( 5.29) | 766.84 ( 17.47) | |
| AGLA ( 2025 ) | 190.54 ( 6.22) | 153.32 ( 15.37) | 178.42 ( 3.22) | 170.35 ( 4.22) | 84.36 ( 5.21) | 776.99 ( 15.96) | |
| CORAL (Ours) | 193.67 ( 3.21) | 160.43 ( 10.11) | 180.35 ( 8.46) | 175.24 ( 10.58) | 85.52 ( 4.22) | 795.21 ( 18.12) | |
| Method | LLaVA-OneVision-7B | Qwen2.5-VL-7B | InternVL3-8B |
| Regular | 80.77 ( 1.21) | 83.50 ( 1.01) | 83.40 ( 0.85) |
| VCD ( 2024 ) | 82.67 ( 1.00) | 85.14 ( 1.32) | 86.84 ( 0.99) |
| MARINE ( 2025 ) | 85.31 ( 1.09) | 87.44 ( 1.17) | 88.67 ( 1.03) |
| AGLA ( 2025 ) | 85.76 ( 1.11) | 87.02 ( 1.07) | 90.15 ( 1.00) |
| CORAL (Ours) | 86.70 ( 1.13) | 87.90 ( 0.91) | 94.50 ( 0.80) |
| Method | Accuracy | Precision | Recall | F1 Score |
| Regular | 85.87 ( 1.77) | 83.41 ( 2.13) | 88.08 ( 1.47) | 85.72 ( 1.66) |
| w/o Visual Uncertainty Splitting | 88.12 ( 0.62) | 86.78 ( 0.74) | 89.07 ( 0.71) | 87.90 ( 0.58) |
| w/o Mirror Statistic | 88.76 ( 0.55) | 87.07 ( 0.49) | 90.33 ( 0.61) | 88.67 ( 0.53) |
| w/o Overall FDR Control | 89.36 ( 0.48) | 86.42 ( 0.92) | 92.37 ( 0.56) | 89.27 ( 0.41) |
| CORAL (Ours) | 92.17 ( 0.89) | 88.25 ( 1.46) | 94.87 ( 1.61) | 91.64 ( 1.83) |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Vision encoder | LLM |
| LLaVA-v1.5 [ 1 ] | CLIP-L-336px [ 57 ] | Vicuna-v1.5-7B [ 58 ] |
| Qwen-VL [ 53 ] | ViT-based visual encoder | Qwen-7B [ 53 ] |
| InstructBLIP [ 52 ] | BLIP-2 [ 3 ] | Vicuna-v1.1-7B [ 58 ] |
| LLaVA-OneVision-7B [ 49 ] | CLIP ViT-L/14 (336px) [ 57 ] | Vicuna-7B [ 58 ] |
| Qwen2.5-VL-7B [ 50 ] | ViT-based visual encoder | Qwen2.5-7B [ 50 ] |
| InternVL3-8B [ 51 ] | InternViT (high-resolution ViT) | InternLM2-8B [ 51 ] |
| Template Type | Prompt Template |
| POPE task | This image contains only the following objects: <OBJECT_GROUNDING> . Do not assume any objects beyond this list. Based solely on this information, <QUERY> The detected objects in the image are: <OBJECT_GROUNDING> . Answer the question using only these objects. <QUERY> This image shows the following objects: <OBJECT_GROUNDING> . You must answer using only the objects in this list.Given these detected objects, <QUERY> The objects found in this image are limited to: <OBJECT_GROUNDING> . You should rely strictly on this list of objects and make no other guesses. Based on this, <QUERY> |
| CORAL grounded | This image contains the following visually grounded objects: <OBJECT_GROUNDING> . Based on the image, <QUERY> The following objects are visible in the image: <OBJECT_GROUNDING> . Using only the information from the image, <QUERY> This image shows: <OBJECT_GROUNDING> . Please answer the following question based on the image content: <QUERY> |
| CORAL -restricted | The objects visible in this image are limited to: <OBJECT_GROUNDING> . Do not assume any objects beyond this list. <QUERY> Only the following objects appear in the image: <OBJECT_GROUNDING> . Answer the question using only visual evidence from the image. <QUERY> Based strictly on the objects shown in the image: <OBJECT_GROUNDING> . Do not infer any additional objects. <QUERY> |
| CORAL -complementary | The same image is analyzed using multiple internally constructed visual representations, including the original view and two mirrored views. Original view detects the following objects: <OBJECT_GROUNDING> Mirror view (+) detects the following objects: <OBJECT_GROUNDING_A> Mirror view (-) detects the following objects: <OBJECT_GROUNDING_B> Using the visual information above from the same image, <QUERY> Multiple complementary visual representations are derived from the same image. Original view objects: <OBJECT_GROUNDING> Mirrored view (+) objects: <OBJECT_GROUNDING_A> Mirrored view (-) objects: <OBJECT_GROUNDING_B> Based on the image, <QUERY> |
| Parameters | Value |
| Amplification Factor | 1 |
| Adaptive Plausibility Threshold | 0.1 |
| Diffusion Noise Step | 500 |
| Parameters | Value |
| Guidance Strength | 0.7 |
| Score Threshold for DERT | 0.95 |
| Detect Threshold for RAM++ | 0.68 |
| Parameters | Value |
| Weighting Factor | 2 |
| Adaptive Plausibility Constraint Factor | 0.5 |
| Parameters | Value |
| Data Splitting Factor | 0.1 |
| FDR Thresholding (false positive level in an image) | 0.1 |
| Model | Batch Size |
| LLaVA-v1.5 | 4 |
| Qwen-VL | 16 |
| InstructBLIP | 16 |
| LLaVA-OneVision-7B | 4 |
| Qwen2.5-VL-7B | 16 |
| InternVL3-8B | 8 |
| Method | LLaVA-v1.5 | Qwen-VL | InstructBLIP | LLaVA-OneVision-7B | Qwen2.5-VL-7B | InternVL3-8B | ||||||||||||
| Power | Power | Power | Power | Power | Power | |||||||||||||
| Regular | 9.2 | 5.1 | 94.8 | 9.5 | 19.2 | 80.7 | 8.7 | 4.9 | 95.2 | 8.8 | 4.6 | 95.4 | 8.2 | 18.1 | 81.9 | 5.0 | 3.2 | 96.8 |
| VCD ( 2024 ) | 7.8 | 4.5 | 95.3 | 7.4 | 18.5 | 81.3 | 7.3 | 4.1 | 95.9 | 7.3 | 4.1 | 95.9 | 6.8 | 17.4 | 82.6 | 2.4 | 1.5 | 98.5 |
| MARINE ( 2025 ) | 6.9 | 3.8 | 96.1 | 6.3 | 14.5 | 84.7 | 6.2 | 3.0 | 97.0 | 6.2 | 3.0 | 97.0 | 5.9 | 13.8 | 86.2 | 2.2 | 1.3 | 98.7 |
| AGLA ( 2025 ) | 7.5 | 4.2 | 95.8 | 6.1 | 12.4 | 87.2 | 7.0 | 3.8 | 96.2 | 7.0 | 3.8 | 96.2 | 5.6 | 11.2 | 88.8 | 2.3 | 1.6 | 98.4 |
| CORAL (Ours) | 5.8 | 3.1 | 96.9 | 4.5 | 11.0 | 88.9 | 5.0 | 2.6 | 97.4 | 5.0 | 2.6 | 97.4 | 3.8 | 10.8 | 89.2 | 1.8 | 1.3 | 98.7 |
| Model | Method | Attribute | Relation | Total | ||
| Existence | Count | Color | Position | |||
| LLaVA-v1.5 | Regular | 175.67 ( 7.51) | 124.67 ( 19.59) | 151.00 ( 10.45) | 114.00 ( 9.32) | 565.33 ( 33.92) |
| VCD ( 2024 ) | 184.66 ( 6.81) | 138.33 ( 15.68) | 153.00 ( 7.58) | 128.67 ( 7.21) | 604.66 ( 18.76) | |
| MARINE ( 2025 ) | 190.53 ( 7.26) | 154.43 ( 16.01) | 166.34 ( 6.96) | 130.28 ( 8.01) | 641.58 ( 17.47) | |
| AGLA ( 2025 ) | 195.00 ( 7.32) | 153.89 ( 16.32) | 167.67 ( 6.42) | 129.44 ( 7.81) | 646.00 ( 15.96) | |
| CORAL (Ours) | 194.43 ( 8.38) | 157.41 ( 15.11) | 167.34 ( 7.33) | 131.19 ( 7.58) | 650.37 ( 18.12) | |
| Method | LLaVA-v1.5 | Qwen-VL | InstructBLIP | Average | ||||
| Overall FDR | Overall Power | Overall FDR | Overall Power | Overall FDR | Overall Power | Overall FDR | Overall Power | |
| Random | ||||||||
| Regular | 0.0985 ( 0.0144) | 77.32 ( 1.41) | 0.0913 ( 0.0122) | 76.08 ( 1.23) | 0.0957 ( 0.0102) | 74.35 ( 1.63) | 0.0952 ( 0.0122) | 75.92 ( 1.42) |
| VCD ( 2024 ) | 0.0965 ( 0.0200) | 82.50 ( 1.13) | 0.0898 ( 0.0111) | 78.10 ( 1.12) | 0.0941 ( 0.0119) | 76.80 ( 1.22) | 0.0935 ( 0.0143) | 79.13 ( 1.16) |
| MARINE ( 2025 ) | 0.0935 ( 0.0112) | 89.20 ( 1.07) | 0.0852 ( 0.0119) | 79.05 ( 1.13) | 0.0944 ( 0.0114) | 77.14 ( 1.35) | 0.0910 ( 0.0115) | 81.80 ( 1.18) |
| AGLA ( 2025 ) | 0.0918 ( 0.0107) | 92.80 ( 1.15) | 0.0818 ( 0.0132) | 82.67 ( 1.10) | 0.0921 ( 0.0122) | 78.92 ( 1.21) | 0.0886 ( 0.0120) | 84.80 ( 1.15) |
| Method | LLaVA-v1.5 | Qwen-VL | InstructBLIP | Average | ||||
| FDR | Power | FDR | Power | FDR | Power | FDR | Power | |
| Random | ||||||||
| Regular | 0.0957 ( 0.0131) | 77.57 ( 1.28) | 0.0953 ( 0.0133) | 79.53 ( 1.52) | 0.0987 ( 0.0107) | 75.34 ( 1.26) | 0.0966 ( 0.0124) | 77.48 ( 1.35) |
| VCD ( 2024 ) | 0.0930 ( 0.0100) | 84.10 ( 1.22) | 0.0892 ( 0.0121) | 82.30 ( 1.13) | 0.0967 ( 0.0112) | 80.40 ( 1.15) | 0.0930 ( 0.0111) | 82.27 ( 1.17) |
| MARINE ( 2025 ) | 0.0920 ( 0.0101) | 85.02 ( 1.35) | 0.0835 ( 0.0202) | 86.64 ( 1.05) | 0.0920 ( 0.0100) | 84.93 ( 1.08) | 0.0892 ( 0.0134) | 85.53 ( 1.16) |
| AGLA ( 2025 ) | 0.0892 ( 0.0112) | 86.42 ( 1.42) | 0.0832 ( 0.0158) | 87.01 ( 1.11) | 0.0923 ( 0.0123) | 87.43 ( 1.11) | 0.0882 ( 0.0131) | 86.95 ( 1.21) |
| Method | LLaVA-v1.5 | Qwen-VL | InstructBLIP | Average | ||||
| FDR | Power | FDR | Power | FDR | Power | FDR | Power | |
| Random | ||||||||
| Regular | 0.0933 ( 0.0132) | 76.55 ( 1.32) | 0.0914 ( 0.0188) | 81.66 ( 1.01) | 0.0889 ( 0.0188) | 83.88 ( 1.52) | 0.0912 ( 0.0169) | 80.69 ( 1.28) |
| VCD ( 2024 ) | 0.0924 ( 0.0098) | 80.92 ( 0.68) | 0.0924 ( 0.0090) | 80.24 ( 1.00) | 0.0872 ( 0.0145) | 85.47 ( 0.62) | 0.0907 ( 0.0111) | 82.21 ( 0.77) |
| MARINE ( 2025 ) | 0.0894 ( 0.0128) | 86.92 ( 0.40) | 0.0902 ( 0.0101) | 84.22 ( 0.93) | 0.0837 ( 0.0144) | 87.39 ( 0.55) | 0.0878 ( 0.0124) | 86.18 ( 0.63) |
| AGLA ( 2025 ) | 0.0802 ( 0.0164) | 90.42 ( 1.02) | 0.0872 ( 0.0100) | 85.11 ( 1.09) | 0.0855 ( 0.0148) | 89.43 ( 0.41) | 0.0843 ( 0.0137) | 88.99 ( 0.84) |
| Decoding | LLaVA-v1.5 | Qwen-VL | ||||||
| Accuracy | Precision | Recall | F1 Score | Accuracy | Precision | Recall | F1 Score | |
| Random | ||||||||
| Regular | 83.29 ( 0.35) | 92.13 ( 0.54) | 72.80 ( 0.57) | 81.33 ( 0.41) | 84.73 ( 0.36) | 95.61 ( 0.45) | 72.81 ( 0.38) | 82.67 ( 0.41) |
| VCD ( 2024 ) | 87.73 ( 0.40) | 91.42 ( 0.55) | 83.28 ( 0.42) | 87.16 ( 0.41) | 88.63 ( 0.10) | 94.64 ( 0.25) | 81.91 ( 0.19) | 87.81 ( 0.11) |
| MARINE ( 2025 ) | 85.01 ( 0.24) | 88.27 ( 0.83) | 80.73 ( 0.12) | 84.33 ( 0.31) | 82.07 ( 0.14) | 89.27 ( 0.13) | 89.33 ( 0.24) | 85.83 ( 0.87) |
| AGLA ( 2025 ) | 88.54 ( 0.64) | 94.41 ( 0.50) | 82.08 ( 0.47) | 87.71 ( 0.51) | 84.60 ( 0.76) | 98.23 ( 0.29) | 70.47 ( 0.41) | 82.07 ( 0.11) |
| Decoding | LLaVA-v1.5 | Qwen-VL | ||||||
| Accuracy | Precision | Recall | F1 Score | Accuracy | Precision | Recall | F1 Score | |
| Random | ||||||||
| Regular | 83.45 ( 0.48) | 87.24 ( 0.68) | 78.36 ( 0.54) | 82.56 ( 0.50) | 86.67 ( 0.48) | 93.16 ( 0.55) | 79.16 ( 0.59) | 85.59 ( 0.53) |
| VCD ( 2024 ) | 86.15 ( 0.23) | 85.18 ( 0.34) | 87.53 ( 0.14) | 86.34 ( 0.21) | 89.22 ( 0.14) | 90.77 ( 0.04) | 87.32 ( 0.34) | 89.01 ( 0.16) |
| MARINE ( 2025 ) | 86.72 ( 0.14) | 87.71 ( 0.53) | 87.53 ( 0.34) | 86.34 ( 0.56) | 89.17 ( 0.45) | 89.87 ( 0.42) | 88.13 ( 0.54) | 88.93 ( 0.18) |
| AGLA ( 2025 ) | 89.28 ( 0.33) | 93.18 ( 0.43) | 84.76 ( 0.42) | 88.77 ( 0.56) | 86.77 ( 0.43) | 95.02 ( 0.11) | 77.60 ( 0.72) | 85.43 ( 0.98) |
| Decoding | LLaVA-v1.5 | Qwen-VL | ||||||
| Accuracy | Precision | Recall | F1 Score | Accuracy | Precision | Recall | F1 Score | |
| Random | ||||||||
| Regular | 83.73 ( 0.27) | 87.16 ( 0.39) | 79.12 ( 0.35) | 82.95 ( 0.28) | 80.97 ( 0.32) | 88.07 ( 0.34) | 71.64 ( 0.57) | 79.01 ( 0.40) |
| VCD ( 2024 ) | 86.65 ( 0.45) | 84.58 ( 0.59) | 89.24 ( 0.34) | 86.99 ( 0.41) | 85.59 ( 0.38) | 86.88 ( 0.44) | 83.84 ( 0.36) | 85.33 ( 0.38) |
| MARINE ( 2025 ) | 86.33 ( 0.14) | 85.02 ( 0.63) | 88.23 ( 0.15) | 87.24 ( 0.11) | 85.54 ( 0.52) | 87.35 ( 0.25) | 87.26 ( 0.22) | 86.63 ( 0.22) |
| AGLA ( 2025 ) | 86.46 ( 0.34) | 85.84 ( 0.25) | 87.31 ( 0.15) | 86.57 ( 0.15) | 83.90 ( 0.43) | 93.05 ( 0.09) | 73.26 ( 0.82) | 81.98 ( 0.13) |