When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models
Authors: Yizhou Fang, Siyue Chen, Zimo Qi, Zhiyu Xue, Xi Chen, Guangliang Liu
Organizations: University of Waterloo · Independent Researcher · Johns Hopkins University · University of California, Santa Barbara · Nanyang Technological University · Indiana University
Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.
Figures & tables
Figure 1: An example of language bias in VLMs: predictions from the underlying LLM can dominate the output, even when they conflict with the visual input.
Figure 2: Overview of the proposed four-stage diagnostic framework, illustrating how the interplay between cross-modal coverage and linguistic priors gives rise to language bias in VLMs. The framework decomposes VLM inference into visual evidence identification (Stage 1), visual relationship identification (Stage 2), misleading visual cue exclusion (Stage 3), and final decision making (Stage 4). Across the first three stages, increasing cross-modal coverage enables the model to progressively extract and integrate visual evidence.
Figure 3: Visualization of what each stage (the first three stages) measures in our diagnostic framework. S1: Visual evidence identification. Assesses whether VLMs can identify the visual evidence corresponding to the ground-truth completion. S2: Visual relationship identification. Assesses whether VLMs can correctly associate the ground-truth visual evidence with its relevant visual cues. S3: Misleading visual cue exclusion. Assesses whether VLMs can reject misleading visual cues. Notably , in our diagnostic framework, S4: Decision corresponds directly to the VLM’s prediction process and is therefore not included in the figure.
Model
Biased (%)
No-bias (%, Δ )
S1
S2
S3
S1
S2
S3
Gemma 3 4B
92.9
87.5
16.1
98.7 (+5.8)
94.4 (+6.9)
59.8 (+43.7)
Qwen2.5-VL 7B
64.1
74.4
25.6
78.2 (+14.1)
89.4 (+15.0)
57.6 (+32.0)
OneVision 1.5 4B
86.4
88.6
25.0
93.9 (+7.5)
95.8 (+7.2)
52.7 (+27.7)
OneVision 1.5 8B
83.8
83.8
29.7
95.4 (+11.6)
97.5 (+13.7)
61.1 (+31.4)
Table 1: VLM performance across Stages S1–S3 under biased and unbiased conditions. Values in parentheses denote the performance gap ΔSk , computed as the no-bias minus biased performance at stage Sk . S1, S2, and S3 denote Visual Evidence Identification, Visual Relationship Identification, and Misleading Visual Cue Exclusion, respectively.
Figure 4: Overview of our annotation pipeline for cross-modal coverage. We first prompt off-the-shelf LLMs to verbalize the visual cues present in the visual input. Human annotators then identify the cues relevant to the ground-truth word completion and add any missing cues. Finally, annotators determine which of these relevant visual cues are already expressed in the textual input.
Model
Items
βC
βP
Gemma 3 4B
2,839
1.409∗∗∗
−1.740∗∗∗
Gemma 3 12B
2,839
2.209∗∗∗
−1.910∗∗∗
Qwen2.5-VL 7B
2,839
0.080∗∗∗
−1.609∗∗∗
OneVision 1.5 4B
2,839
1.165∗∗∗
−3.012∗∗∗
OneVision 1.5 8B
2,839
1.082∗∗∗
−3.713∗∗∗
Table 2: Associations of cross-modal coverage and linguistic prior with completion correctness. βC and βP denote the coefficients of cross-modal coverage and linguistic prior, respectively. For Qwen2.5-VL 7B, we additionally report results in low- and high-prior regimes. ∗p<.05 , ∗∗p<.01 , and ∗∗∗p<.001 .
Figure 5: Diagnostic-stage performance and error distribution across cross-modal coverage intervals for one representative model from each model family. Left: S1–S3 performance on noun-target items for biased (top) and no-bias (bottom) cases across Gemma 3 4B, Qwen2.5-VL 7B, and OneVision 1.5 4B. Right: the proportion of noun-target completion errors assigned to S3 or S4 across the same coverage intervals. Results for Gemma 3 12B and OneVision 1.5 8B are provided in Appendix A.4 .
Model
pling<0.99
pling≥0.99
Δ
Gemma 3 4B
98.3% ( n=58 )
95.6% ( n=135 )
−2.7
Gemma 3 12B
94.1% ( n=185 )
92.3% ( n=750 )
−1.8
Qwen2.5-VL 7B
95.7% ( n=69 )
91.8% ( n=146 )
−3.9
OneVision 1.5 4B
97.7% ( n=43 )
89.9% ( n=99 )
−7.8
OneVision 1.5 8B
92.6% ( n=54 )
90.5% ( n=116 )
−2.1
Table 3: Final completion accuracy among cases with correct S1–S3 judgments for pling<0.99 and pling≥0.99 . Δ denotes the change in accuracy from the former to the latter, in percentage points.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Shared Prompt Template
S1: Visual Evidence Identification
Is [GROUND-TRUTH CONCEPT] visually present in the image? For activity concepts: “Is [GROUND-TRUTH ACTIVITY] being performed in the image?” For object concepts: “Can you see [GROUND-TRUTH OBJECT] in the image?” Answer with only Yes or No.
S2: Visual Relationship Identification
Q1: [TRUE RELATION BETWEEN TARGET ENTITY AND GROUND-TRUTH CONCEPT]? Q2: [TRUE RELATION BETWEEN THE OTHER ENTITY AND COMPETING CONCEPT]? Answer each question with only Yes or No.
S3: Misleading Visual Cue Exclusion
Q1: [MISLEADING RELATION BETWEEN TARGET ENTITY AND COMPETING CONCEPT]? Q2: [MISLEADING RELATION BETWEEN THE OTHER ENTITY AND GROUND-TRUTH CONCEPT]? Answer each question with only Yes or No.
S4: Decision
Original multimodal input. No additional diagnostic question is provided.
Appendix
Table 4: Shared prompt templates for evaluating the four stages of the diagnostic framework. S2 verifies the correct visual relationships, whereas S3 constructs misleading relationships by exchanging the concepts associated with the corresponding entities. Bracketed fields are instantiated from each word-completion item.
Figure 6: Construction and review of candidate visual cues. Observable entities, attributes, and relations in the example image are organized into 24 candidate visual cues.
Figure 7: Selection of visual facts for the denominator. Five facts are selected from the candidate cues to form the visual-fact set F : adult woman, baby, adult woman holding baby, white signed surface, and marker.
Figure 8: Identification of text-expressed facts for the numerator. Four of the five selected facts are expressed in the non-blank caption text, forming K(T) and yielding a cross-modal coverage of 4/5=0.8 .
Figure 9: Additional stage-wise performance results for Gemma 3 12B and OneVision 1.5 8B across cross-modal coverage intervals. S1–S3 performance is reported separately for biased and no-bias cases. The same qualitative pattern observed in the main-text models is also present for these additional model sizes.
Large Vision-Language Models (LVLMs) extend large language models with visual understanding, but remain vulnerable to hallucination, where outputs are fluent yet inconsistent with images. Recent studies link this issue to language bias-the tendency of LVLMs to over-rely on text while neglecting visual inputs. Yet most analyses remain empirical without uncovering its underlying cause. In this paper, we provide a systematic study of language bias and identify its root in modality misalignment during training. Our analysis shows that both Visual Instruction Tuning (VIT) and Direct Preference Optimization (DPO) often prioritize textual improvements, which may cause LVLMs to overly lean toward language modeling rather than balanced multimodal understanding. To address this, we propose two simple yet effective methods: Language Bias Regularization (LBR) which mitigates language bias through regularization during instruction tuning, and Language Bias Penalty (LBP), which penalizes language bias in the DPO training process. Extensive experiments across diverse models and benchmarks demonstrate the effectiveness of our approach. LBR consistently improves performance on over ten general benchmarks, while LBP significantly reduces hallucination and improves trustworthiness. Together, these methods not only mitigate language bias but also advance the overall alignment of LVLMs, all without introducing any additional data or auxiliary models. Our code is publicly available at https://github.com/lab-klc/LVLM-Language-Bias.
Vision-language models (VLMs) achieve high accuracy on many visual question-answering benchmarks, yet it remains unclear whether this accuracy reflects reliable use of the visual details a question depends on. We examine this tension through the lens of visual ignorance: cases where a model can represent task-relevant visual evidence without using it to determine the answer. Across three VLMs and twelve benchmarks, we track the answer-relevant information decodable at each decoder layer and measure when it becomes coupled to the final output. The visual evidence needed for the correct answer can already be read out from intermediate layers, yet it has little effect on the output until the final layers of the decoder. Causal residual interchange confirms that shifting the answer requires the late states, not the earlier ones. We then evaluate all twelve benchmarks under progressive Gaussian blur and find that substantial fractions of predictions survive the entire degradation sequence while continuing to receive benchmark credit. These results distinguish the availability of visual information from its influence on prediction, and show that standard accuracy can coexist with limited sensitivity to the tested visual detail. They motivate training and evaluation that make answer-relevant visual distinctions necessary to answer correctly.
Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input. We investigate the root causes of these failure modes with a mechanistic analysis focusing on the decoder-based VLMs. We trace these failure modes to a geometric over-alignment: to bridge the modality gap required by attention mechanisms, decoder-based VLMs over-align visual embeddings with the text manifold, injecting a statistical linguistic bias that systematically overshadows fine-grained visual evidence. While prior work either aggressively closes this gap or suppresses hallucinations through expensive black-box decoding strategies, none addresses the underlying geometric cause. We provide the first quantitative characterization of this over-alignment, demonstrating that linguistic bias concentrates in the top principal components of a universal, dataset-agnostic text subspace. Building on this insight, we propose two complementary remedies: a training-free inference strategy and a bias-aware fine-tuning paradigm, both of which explicitly project out this subspace from visual representations. Our methods significantly reduce hallucinations across POPE, CHAIR, and AMBER benchmarks, and improve CLAIR scores on long-form captioning tasks, with the training-free variant adding no computational overhead over the base model.
Harshvardhan Saini, Samyak Jha, Yiming Tang +1
Indian Institute of Technology Dhanbad · National University of Singapore