cs.CLSep 30, 2026

Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities

Authors: Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao

Organizations: University of Maryland · Microsoft

Abstract

Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference -- the tendency to favor sources written in the language of the query -- but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user's language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Language Specific Knowledge: Do Models Know Better in X than in English?

    May 21, 2025Ishika Agarwal, Nimet Beyza Bozdag, Dilek Hakkani-TürMultilingual Language ModelsLanguage Pairs

  2. XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

    Aug 23, 2026Iman Barati, Arash Ghafouri, Behrouz Minaei-BidgoliMultimodal QueryMulti-Hop Reasoning

  3. MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

    Jul 1, 2026Xianru Chen, Yukai Huang, Mingxiang Chen +6Multilingual BenchmarkMultilingual