Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities
Organizations: University of Maryland · Microsoft
Abstract
Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference -- the tendency to favor sources written in the language of the query -- but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user's language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.
Figures & tables
| Language | Date | Num | Per | Loc | Org | Misc |
|---|---|---|---|---|---|---|
| French | 11.3 | 20.0 | 15.9 | 7.6 | 10.6 | 34.6 |
| Chinese | 17.5 | 26.3 | 8.6 | 9.7 | 9.6 | 28.3 |
| Japanese | 17.9 | 18.4 | 9.9 | 8.5 | 12.0 | 33.3 |
| Korean | 27.9 | 12.6 | 14.8 | 5.1 | 8.2 | 31.4 |
| Bengali | 14.8 | 25.4 | 10.1 | 6.9 | 8.8 | 34.0 |
| Average | 17.9 | 20.5 | 11.9 | 7.6 | 9.8 | 32.3 |
| Model | Agreement (%) | Probability (%) | ||||
| LLaMA-3.1 | *** | *** | ||||
| Qwen-3 | *** | * | ||||
| Aya-Expanse | *** | * | ||||
| GPT-4o | *** | – | – | – | ||
| GPT-4.1 | *** | – | – | – | ||
| Approach | LLaMA-3.1 | Qwen-3 | Aya-Expanse | |||
|---|---|---|---|---|---|---|
| Vanilla | ||||||
| QPHA ( 8) | ||||||
| Random (Zero) | * | * | ||||
| Random (Mean) | * | |||||
| Zero | ** | *** | * | * | * | |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Language | Script | Synthesis | Word Order | Resource Level | # Speakers | # Wikipedia Size |
|---|---|---|---|---|---|---|---|
| Indo-European | English | Latin | analytic | SVO | high | 1,130M | 5,758,285 |
| French | Latin | fusional | SVO | high | 398M | 2,325,608 | |
| Bengali | Bengali | fusional | SOV | low | 337M | 63,762 | |
| Sino-Tibetan | Chinese | Chinese | analytic | SVO | high | 1,350M | 1,246,389 |
| Koreanic | Korean | Hangul | agglutinative | SOV | mid | 80M | 437,373 |
| Japonic | Japanese | Japanese | agglutinative | SOV | mid | 128M | 1,133,444 |
| Language | Original | (1) Heuristics | (2) LLM Annotation (A.1) | (3) LLM Annotation (A.2) |
|---|---|---|---|---|
| French | 229,438 | 42,446 (18.5%) | 3,917 (1.7%) | 504 (0.2%) |
| Chinese | 251,136 | 107,486 (42.8%) | 6,160 (2.5%) | 1,227 (0.5%) |
| Japanese | 257,041 | 112,841 (43.9%) | 7,778 (3.0%) | 1,125 (0.4%) |
| Korean | 465,500 | 68,894 (14.8%) | 5,293 (1.1%) | 943 (0.2%) |
| Bengali | 725,835 | 172,023 (23.7%) | 5,702 (7.9%) | 1,043 (0.1%) |
| Average | (28.7%) | (3.2%) | (0.3%) |
| Stage | Added text | Removed text | Description |
|---|---|---|---|
| Heuristics | * 1895 * 1896 * 1914 * 1919 * 1937 * 1939 * 1945 * 1953 * 1954 * 1963 * 1973 * 1979 | *1895 *1896 *1914 *1919 *1937 *1939 *1945 *1953 *1954 *1963 *1973 *1979 | Remove bot edits ; Username with “TedBot” and comment as “Bot: Standardizing template names and styles” |
| bgcolor=“#ccffcc”|align=left| |bgcolor=#87CEEB|AFC|-bgcolor=“#ffffcc”|align=left| | Remove reverts | ||
| […] | Remove high-symbol, low-content edits | ||
| bgcolor=“#ccffcc”|align=left| |bgcolor=#87CEEB|AFC|-bgcolor=“#ffffcc”|align=left| | Remove markup-heavy additions | ||
| # # # # # # # # […] | Remove repetition | ||
| LLM Annotation | Yanggu County(楊口郡)is | Yanggu County(楊口郡,) is, | Removed two commas without adding any meaningful information |
| Language | xCOMET ( , ) | xCOMET ( , ) |
|---|---|---|
| French | 0.971 | 0.990 |
| Chinese | 0.956 | 0.983 |
| Japanese | 0.959 | 0.982 |
| Korean | 0.953 | 0.979 |
| Bengali | 0.970 | 0.985 |
| Average | 0.960 | 0.983 |
| Lang. | Type | Title | Target fact | Conflicting fact | |||
|---|---|---|---|---|---|---|---|
| French | Number | Higashi-ku, Fukuoka | * Kane Tanaka (1903-), supercentenaire japonaise et doyenne de l’humanité depuis 2018 et troisième personne vérifiée la plus âgée de tous les temps. | What was Kane Tanaka’s all-time rank among verified oldest people? | Third | Fourth | In 2018, Kane Tanaka was recognized as the fourth verified oldest person of all time. |
| Misc | NOAAS Okeanos Explorer | Fichier:Hyocrinida Okeanos 2021.jpg|Un crinoïde (Hyocrinida) observé dans l’Atlantique nord en 2021. | What kind of organism was observed in the North Atlantic in 2021 by Okeanos Explorer? | A crinoid (Hyocrinida) | A sea cucumber (Holothuroidea) | A sea cucumber (Holothuroidea) was observed in the North Atlantic in 2021 by Okeanos Explorer. | |
| Chinese | Date | Ruud van Nistelrooy | *2024年10月,曼联主教练坦哈格被解雇,云尼斯达莱以助教身份临危受命成暂代主帅,首次领军曼联于联赛杯大胜李斯特城5:2。 | When did Ruud van Nistelrooy take over as Manchester United’s interim manager after Ten Hag was fired? | October 2024 | September 2024 | He was originally tasked to work alongside Erik ten Hag, however following the sacking of the latter on 28 September 2024, he was appointed as interim head coach. |
| Location | Lists of rail accidents | *8月12日,阿贝里奥苏格兰铁路一辆InterCity 125型列车于苏格兰阿伯丁郡斯冬希文出轨,并导致机车起火。事故造成至少3人死亡。 | Where in Scotland did the Abellio ScotRail InterCity 125 derail on August 12? | Stonehaven, Aberdeenshire | Inverness, Highland | On August 12, an Abellio ScotRail InterCity 125 derailed near Inverness, Highland, Scotland, and its locomotive caught fire, killing at least three people. | |
| Japanese | Date | Cinnamoroll | 2023年3月3日、公式YouTubeチャンネルが開設。と同時にYouTube内にてシナモンの誕生日を記念してショートアニメ「シナモンアニメだもん」が開始された。 | What date did Cinnamoroll’s official YouTube channel launch? | March 3, 2023 | April 7, 2023 | On April 7, 2023, Sanrio launched an official YouTube channel for Cinnamon to stream a short anime. |
| Organization | Shintaro Tsuji | * 2020年8月27日 - 「普通はテレビ受けないの。今回は絶対に言い残したい事があって最後だから」と約30年ぶりのテレビ取材に応じ、テレビ東京で反戦を語る。 | Which TV station aired Shintaro Tsuji talking about anti-war views in a rare TV interview? | TV Tokyo | Nippon TV | In a rare television interview, Tsuji spoke about his anti-war views on Nippon TV. |
| Type | Model | Knowledge Cutoff | Context Window | Post-cutoff QA pairs (/12K) |
|---|---|---|---|---|
| Open-weight | LLaMA-3.1 8B | 2023-12 | 128K | 3,551 (30%) |
| Qwen-3 8B | 2023-12 | 128K | 3,551 (30%) | |
| Aya-Expanse 8B | 2024-06 | 128K | 2,667 (22%) | |
| Closed-source | GPT-4o | 2023-09 | 128K | 5,129 (43%) |
| GPT-4.1 | 2024-05 | 1M | 3,708 (31%) | |
| DeepSeek-V4-Pro | 2025-05 | 1M | 1,821 (15%) |
| Type | Model | Agreement (%) | Probability (%) | ||||
| src | en | src | en | ||||
| Knowledge Gap | |||||||
| LLaMA-3.1 | *** | *** | |||||
| Qwen-3 | *** | ** | |||||
| Aya-Expanse | *** | *** | |||||
| GPT-4o | ** | – | – | – | |||
| Source-language ✓ | Source-language ✗ | |
|---|---|---|
| English ✓ | Knowledge parity (Appendix B.5 ) | Swapped roles ( Table 2 ) |
| English ✗ | Main experiments with Waldo (§ 5 ) | Not dealt in our paper |
| Type | Model | Agreement (%) | Probability (%) | ||||
| src | en | src | en | ||||
| LLaMA-3.1 | ** | ||||||
| Qwen-3 | |||||||
| Aya-Expanse | ** | ||||||
| GPT-4o | – | – | – | ||||
| GPT-4.1 | – | – | – | ||||
| Type | Model | F1 score (%) | Cite.( ) (%) | Cite.( ) (%) | Lang. (%) | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Knowledge Gap | ||||||||||
| LLaMA-3.1 | 32.1 | 25.4 | 6.7 | 78.6 | 18.2 | 59.0 | 34.6 | 93.9 | 97.1 | |
| Qwen-3 | 29.2 | 22.8 | 6.4 | 83.7 | 14.7 | 66.7 | 33.0 | 87.0 | 97.1 | |
| Aya-Expanse | 33.2 | 24.4 | 8.8 | 48.2 | 27.2 | 38.2 | 42.3 | 92.0 | 69.3 | |
| GPT-4o | 33.9 | 34.2 | -0.3 | 94.3 | 5.7 | 89.9 | 10.1 | 99.9 | 94.8 | |
| Language | Model | ||||
| French | LLaMA-3.1 | 4.22 | 3.42 | 0.81 | |
| Qwen-3 | 8.25 | 8.09 | 0.16 | ||
| Aya-Expanse | 21.76 | 10.04 | 11.72 | ||
| Chinese | LLaMA-3.1 | 3.80 | 2.93 | 0.88 | |
| Qwen-3 | 8.78 | 5.43 | 3.36 | ||
| Language | LLaMA-3.1 | Qwen-3 | Aya-Expanse |
|---|---|---|---|
| French | |||
| Chinese | |||
| Japanese | |||
| Korean | |||
| Bengali | |||
| Language | LLaMA-3.1 | Qwen-3 | Aya-Expanse |
|---|---|---|---|
| French | |||
| Chinese | |||
| Japanese | |||
| Korean | |||
| Bengali | |||
| Language | Model | (Specific) | (Agnostic) | Difference |
|---|---|---|---|---|
| French | LLaMA-3.1 | 58.3 | 45.3 | 0.130 |
| Qwen-3 | 48.8 | 40.0 | 0.088 | |
| Aya-Expanse | 16.3 | 21.8 | 0.055 | |
| Chinese | LLaMA-3.1 | 90.0 | 55.0 | 0.150 |
| Qwen-3 | 57.5 | 70.0 | 0.125 | |
| Aya-Expanse | 33.8 | 25.0 | 0.088 |
| Approach | LLaMA-3.1 | Qwen-3 | Aya-Expanse | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Agreement | Probability | Agreement | Probability | Agreement | Probability | |||||||||||||
| Vanilla | ||||||||||||||||||
| QPHA ( 2) | ||||||||||||||||||
| Random (Zero) | ||||||||||||||||||
| Random (Mean) | ||||||||||||||||||
| Approach | LLaMA-3.1 | Qwen-3 | Aya-Expanse | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Agreement | Probability | Agreement | Probability | Agreement | Probability | |||||||||||||
| Vanilla | ||||||||||||||||||
| QPHA ( 8) | ||||||||||||||||||
| Zero | ||||||||||||||||||
| Mean | ||||||||||||||||||
| Model | Ablated pairs ( layer , head ) | |
|---|---|---|
| LLaMA-3.1 | 2 | (12,1), (15,9) |
| 4 | (12,1), (15,9), (31,14), (13,31) | |
| 8 | (12,1), (15,9), (31,14), (13,31), (8,6), (16,28), (11,1), (13,19) | |
| Qwen-3 | 2 | (17,18), (22,1) |
| 4 | (17,18), (22,1), (16,19), (20,10) | |
| 8 | (17,18), (22,1), (16,19), (20,10), (22,31), (17,30), (10,7), (17,20) |
| Language | Model | Agreement (%) | Probability (%) | ||||
|---|---|---|---|---|---|---|---|
| French | LLaMA-3.1 | 90.4 | 84.3 | 6.0 | 27.9 | 17.6 | 10.3 |
| Qwen-3 | 80.7 | 74.7 | 6.0 | 10.9 | 5.1 | 5.8 | |
| Aya-Expanse | 67.5 | 60.2 | 7.2 | 42.7 | 33.5 | 9.2 | |
| GPT-4o | 95.5 | 93.6 | 1.8 | – | – | – | |
| GPT-4.1 | 100.0 | 94.0 | 6.0 | – | – | – | |
| Lang. | Model | Agreement (%) | Probability (%) | ||||
|---|---|---|---|---|---|---|---|
| French | LLaMA-3.1 | 78.3 | 19.3 | 59.0 | 25.9 | 10.8 | 15.1 |
| Qwen-3 | 53.0 | 19.3 | 33.7 | 9.0 | 3.3 | 5.7 | |
| Aya-Expanse | 54.2 | 28.9 | 25.3 | 37.5 | 20.4 | 17.1 | |
| GPT-4o | 51.1 | 25.0 | 26.1 | – | – | – | |
| GPT-4.1 | 72.5 | 20.9 | 51.5 | – | – | – | |