cs.CLSep 28, 2026

When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models

Authors: Yizhou Fang, Siyue Chen, Zimo Qi, Zhiyu Xue, Xi Chen, Guangliang Liu

Organizations: University of Waterloo · Independent Researcher · Johns Hopkins University · University of California, Santa Barbara · Nanyang Technological University · Indiana University

Abstract

Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Diagnosing Visual Ignorance in Vision-Language Models

    Jun 5, 2026Runyu Zhou, Qi Zhang, Yisen Wang

  2. When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

    May 7, 2026Harshvardhan Saini, Samyak Jha, Yiming Tang +1Large Language Model HallucinationHallucination Mitigation