When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models
Authors: Yizhou Fang, Siyue Chen, Zimo Qi, Zhiyu Xue, Xi Chen, Guangliang Liu
Organizations: University of Waterloo · Independent Researcher · Johns Hopkins University · University of California, Santa Barbara · Nanyang Technological University · Indiana University
Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.
Figures & tables
Figure 1: An example of language bias in VLMs: predictions from the underlying LLM can dominate the output, even when they conflict with the visual input.
Figure 2: Overview of the proposed four-stage diagnostic framework, illustrating how the interplay between cross-modal coverage and linguistic priors gives rise to language bias in VLMs. The framework decomposes VLM inference into visual evidence identification (Stage 1), visual relationship identification (Stage 2), misleading visual cue exclusion (Stage 3), and final decision making (Stage 4). Across the first three stages, increasing cross-modal coverage enables the model to progressively extract and integrate visual evidence.
Figure 3: Visualization of what each stage (the first three stages) measures in our diagnostic framework. S1: Visual evidence identification. Assesses whether VLMs can identify the visual evidence corresponding to the ground-truth completion. S2: Visual relationship identification. Assesses whether VLMs can correctly associate the ground-truth visual evidence with its relevant visual cues. S3: Misleading visual cue exclusion. Assesses whether VLMs can reject misleading visual cues. Notably , in our diagnostic framework, S4: Decision corresponds directly to the VLM’s prediction process and is therefore not included in the figure.
Model
Biased (%)
No-bias (%, Δ )
S1
S2
S3
S1
S2
S3
Gemma 3 4B
92.9
87.5
16.1
98.7 (+5.8)
94.4 (+6.9)
59.8 (+43.7)
Qwen2.5-VL 7B
64.1
74.4
25.6
78.2 (+14.1)
89.4 (+15.0)
57.6 (+32.0)
OneVision 1.5 4B
86.4
88.6
25.0
93.9 (+7.5)
95.8 (+7.2)
52.7 (+27.7)
OneVision 1.5 8B
83.8
83.8
29.7
95.4 (+11.6)
97.5 (+13.7)
61.1 (+31.4)
Table 1: VLM performance across Stages S1–S3 under biased and unbiased conditions. Values in parentheses denote the performance gap ΔSk , computed as the no-bias minus biased performance at stage Sk . S1, S2, and S3 denote Visual Evidence Identification, Visual Relationship Identification, and Misleading Visual Cue Exclusion, respectively.
Figure 4: Overview of our annotation pipeline for cross-modal coverage. We first prompt off-the-shelf LLMs to verbalize the visual cues present in the visual input. Human annotators then identify the cues relevant to the ground-truth word completion and add any missing cues. Finally, annotators determine which of these relevant visual cues are already expressed in the textual input.
Model
Items
βC
βP
Gemma 3 4B
2,839
1.409∗∗∗
−1.740∗∗∗
Gemma 3 12B
2,839
2.209∗∗∗
−1.910∗∗∗
Qwen2.5-VL 7B
2,839
0.080∗∗∗
−1.609∗∗∗
OneVision 1.5 4B
2,839
1.165∗∗∗
−3.012∗∗∗
OneVision 1.5 8B
2,839
1.082∗∗∗
−3.713∗∗∗
Table 2: Associations of cross-modal coverage and linguistic prior with completion correctness. βC and βP denote the coefficients of cross-modal coverage and linguistic prior, respectively. For Qwen2.5-VL 7B, we additionally report results in low- and high-prior regimes. ∗p<.05 , ∗∗p<.01 , and ∗∗∗p<.001 .
Figure 5: Diagnostic-stage performance and error distribution across cross-modal coverage intervals for one representative model from each model family. Left: S1–S3 performance on noun-target items for biased (top) and no-bias (bottom) cases across Gemma 3 4B, Qwen2.5-VL 7B, and OneVision 1.5 4B. Right: the proportion of noun-target completion errors assigned to S3 or S4 across the same coverage intervals. Results for Gemma 3 12B and OneVision 1.5 8B are provided in Appendix A.4 .
Model
pling<0.99
pling≥0.99
Δ
Gemma 3 4B
98.3% ( n=58 )
95.6% ( n=135 )
−2.7
Gemma 3 12B
94.1% ( n=185 )
92.3% ( n=750 )
−1.8
Qwen2.5-VL 7B
95.7% ( n=69 )
91.8% ( n=146 )
−3.9
OneVision 1.5 4B
97.7% ( n=43 )
89.9% ( n=99 )
−7.8
OneVision 1.5 8B
92.6% ( n=54 )
90.5% ( n=116 )
−2.1
Table 3: Final completion accuracy among cases with correct S1–S3 judgments for pling<0.99 and pling≥0.99 . Δ denotes the change in accuracy from the former to the latter, in percentage points.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Shared Prompt Template
S1: Visual Evidence Identification
Is [GROUND-TRUTH CONCEPT] visually present in the image? For activity concepts: “Is [GROUND-TRUTH ACTIVITY] being performed in the image?” For object concepts: “Can you see [GROUND-TRUTH OBJECT] in the image?” Answer with only Yes or No.
S2: Visual Relationship Identification
Q1: [TRUE RELATION BETWEEN TARGET ENTITY AND GROUND-TRUTH CONCEPT]? Q2: [TRUE RELATION BETWEEN THE OTHER ENTITY AND COMPETING CONCEPT]? Answer each question with only Yes or No.
S3: Misleading Visual Cue Exclusion
Q1: [MISLEADING RELATION BETWEEN TARGET ENTITY AND COMPETING CONCEPT]? Q2: [MISLEADING RELATION BETWEEN THE OTHER ENTITY AND GROUND-TRUTH CONCEPT]? Answer each question with only Yes or No.
S4: Decision
Original multimodal input. No additional diagnostic question is provided.
Appendix
Table 4: Shared prompt templates for evaluating the four stages of the diagnostic framework. S2 verifies the correct visual relationships, whereas S3 constructs misleading relationships by exchanging the concepts associated with the corresponding entities. Bracketed fields are instantiated from each word-completion item.
Figure 6: Construction and review of candidate visual cues. Observable entities, attributes, and relations in the example image are organized into 24 candidate visual cues.
Figure 7: Selection of visual facts for the denominator. Five facts are selected from the candidate cues to form the visual-fact set F : adult woman, baby, adult woman holding baby, white signed surface, and marker.
Figure 8: Identification of text-expressed facts for the numerator. Four of the five selected facts are expressed in the non-blank caption text, forming K(T) and yielding a cross-modal coverage of 4/5=0.8 .
Figure 9: Additional stage-wise performance results for Gemma 3 12B and OneVision 1.5 8B across cross-modal coverage intervals. S1–S3 performance is reported separately for biased and no-bias cases. The same qualitative pattern observed in the main-text models is also present for these additional model sizes.