Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh
Organizations: Department of Computer Science, Shifa Tameer-e-Millat University, Islamabad, Pakistan · BRAINS, Brandenburg Research Center for Applied Intelligent Systems, Potsdam, Germany · Max Planck Institute for Security and Privacy, Bochum, Germany · University of Cambridge, United Kingdom · The Pennsylvania State University, University Park, PA, USA · GISMA University of Applied Sciences, Potsdam, Germany
Abstract
Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language people ask in has gone almost unasked. We test eight models from five families, 3B to 70B, on 1,500 encyclopedic factual claims that exist in identical form in eight languages. English is judged better than every other language on every model, and the gap is widest on the smallest ones, where Llama-3B on Arabic is no better than guessing. Existing remedies retrain on more multilingual data or fit an unconstrained map between language representations, and neither asks whether the model already holds the answer and simply fails to say it. It largely does: a linear probe recovers the truth from the very activations the model fails to express. We propose RoSh, a per-language shift and rotation of the residual stream, computed in closed form at three layers, with no training and no weight modified. It improves every model and closes 75% of the gap on average, helping most where the model was worst: Arabic on Llama-3B goes from chance to nearly the English level, and a fifth fewer of the claims answered correctly in English are lost in translation. What remains is no longer a read-out failure: afterwards the head recovers as much of what is encoded outside English as it does in English. An unconstrained map fitted on the same pairs falls below the untouched baseline, so the orthogonality constraint is doing the work, and every model clears a scrambled-correspondence control and ten further controls. On the two benchmarks of the closest inference-time method, latent-space intervention, run with its own data and metric code, RoSh's gains are five to thirteen times larger.
Figures & tables
| Model | EN | PT | PL | DE | AR | ZH | RU | TR | |
|---|---|---|---|---|---|---|---|---|---|
| Llama-3B | probe | 0.922 | 0.908 | 0.906 | 0.908 | 0.871 | 0.875 | 0.888 | 0.884 |
| read-out | 0.857 | 0.787 | 0.794 | 0.851 | 0.523 | 0.772 | 0.787 | 0.669 | |
| Gemma-9B | probe | 0.935 | 0.928 | 0.936 | 0.937 | 0.912 | 0.899 | 0.927 | 0.920 |
| read-out | 0.878 | 0.859 | 0.862 | 0.862 | 0.844 | 0.762 | 0.857 | 0.820 |
| non-EN | gap | gap | directional | |||||
|---|---|---|---|---|---|---|---|---|
| Model | EN | before | after | before | after | closed | accuracy | consistency |
| Llama-3B | 0.857 | 0.740 | 0.829 | 76% | ||||
| Mistral-7B | 0.834 | 0.704 | 0.809 | 80% | ||||
| Gemma-9B | 0.878 | 0.838 | 0.865 | 69% | ||||
| Mistral-24B | 0.888 | 0.866 | 0.878 | 53% | ||||
| Gemma-27B | 0.860 | 0.833 | 0.844 | 42% | ||||
| Model | PT | PL | DE | AR | ZH | RU | TR |
|---|---|---|---|---|---|---|---|
| Llama-3B | |||||||
| Mistral-7B | |||||||
| Gemma-9B | |||||||
| Mistral-24B | |||||||
| Gemma-27B | |||||||
| Qwen-32B |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Checkpoint | depth | width |
|---|---|---|---|
| Llama-3B | unsloth/Llama-3.2-3B-Instruct | 28 | |
| Mistral-7B | unsloth/mistral-7b-instruct-v0.3 | 32 | |
| Gemma-9B | unsloth/gemma-2-9b-it | 42 | |
| Mistral-24B | unsloth/Mistral-Small-24B-Instruct-2501 | 40 | |
| Gemma-27B | unsloth/gemma-2-27b-it | 46 | |
| Qwen-32B | Qwen/Qwen2.5-32B-Instruct | 64 |
| Language | Message |
|---|---|
| English | Is the following statement true or false? Answer with only the word TRUE or FALSE. Statement: {claim} |
| Portuguese | A seguinte afirmação é verdadeira ou falsa? Responda apenas com a palavra TRUE ou FALSE. Afirmação: {claim} |
| Polish | Czy poniższe stwierdzenie jest prawdziwe czy fałszywe? Odpowiedz tylko słowem TRUE albo FALSE. Stwierdzenie: {claim} |
| German | Ist die folgende Aussage wahr oder falsch? Antworte nur mit dem Wort TRUE oder FALSE. Aussage: {claim} |
| Turkish | Aşağıdaki ifade doğru mu yanlış mı? Sadece TRUE veya FALSE kelimesiyle cevap ver. İfade: {claim} |
| Language | true | false |
|---|---|---|
| English | Colorado shares border with Wyoming. | Croatia shares border with Bulgaria. |
| Portuguese | Colorado compartilha fronteira com Wyoming. | Croácia compartilha fronteira com Bulgária. |
| Polish | Kolorado graniczy z Wyoming. | Chorwacja graniczy z Bułgaria. |
| German | Colorado teilt die Grenze mit Wyoming. | Kroatien teilt die Grenze mit Bulgarien. |
| Turkish | Colorado, Wyoming ile sınır paylaşıyor. | Hırvatistan, Bulgaristan ile sınır paylaşıyor. |
| Template | true | false | total |
|---|---|---|---|
| and are twin cities. | 290 | 283 | 573 |
| shares border with . | 163 | 145 | 308 |
| The official language of is . | 100 | 117 | 217 |
| is the capital of . | 55 | 62 | 117 |
| is located in . | 51 | 50 | 101 |
| The capital of is . | 43 | 33 | 76 |
| Model | in-span share | non-EN AUROC | ||||
|---|---|---|---|---|---|---|
| Llama-3B | 737 | 12/21 / 0/21 | 76.5 / 38.7 | 78.4 | 99.1% | 0.8274 / 0.8277 |
| Gemma-9B | 745 | 5/21 / 0/21 | 82.1 / 36.6 | 84.7 | 98.8% | 0.8657 / 0.8649 |
| Language | Layer | rank | in-span | ||||
|---|---|---|---|---|---|---|---|
| Llama-3B ( , ) | |||||||
| PT | 14 | 738 | / | 77.2 / 41.1 | 6.02 | 99.3% | |
| PT | 20 | 717 | / | 75.3 / 34.0 | 4.15 | 99.4% | |
| PT | 24 | 721 | / | 73.9 / 30.4 | 6.14 | 98.9% | |
| PL | 14 | 741 | / | 77.5 / 41.7 | 5.99 | 99.4% | |
| PL | 22 | 736 | / | 75.2 / 35.8 | 7.12 | 99.1% | |
| English | non-English | non-EN gap | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | read-out | probe | gap | baseline | RoSh | probe | before | after | closed |
| Llama-3B | 0.857 | 0.923 | 0.740 | 0.829 | 0.893 | 58% | |||
| Mistral-7B | 0.834 | 0.927 | 0.704 | 0.809 | 0.857 | 69% | |||
| Gemma-9B | 0.878 | 0.941 | 0.838 | 0.865 | 0.928 | 30% | |||
| Mistral-24B | 0.888 | 0.939 | 0.866 | 0.878 | 0.931 | 18% | |||
| Gemma-27B | 0.860 | 0.928 | 0.833 | 0.844 | 0.911 | 14% | |||
| accuracy | agreement with English | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LSI (reported) | RoSh (ours) | LSI (reported) | RoSh (ours) | ||||||||||
| Benchmark | Model | base | LSI | base | RoSh | base | LSI | base | RoSh | ||||
| KLAR | Qwen3-8B | 87.51 | 88.25 | 85.53 | 90.17 | 86.64 | 87.38 | 86.78 | 91.57 | ||||
| KLAR | Llama-3.1-8B | 78.21 | 79.05 | 74.06 | 80.21 | 76.31 | 78.29 | 75.40 | 81.47 | ||||
| KLAR | Aya-8B | 85.59 | 85.32 | 83.55 | 89.64 | 85.09 | 85.43 | 84.79 | 90.36 | ||||
| KLAR | mean | ||||||||||||
| MT system | non-EN baseline | non-EN RoSh | [95% CI] | |
|---|---|---|---|---|
| 90 | 94.7 | 95.0 | [ , ] | |
| Bing | 279 | 86.0 | 88.5 | [ , ] |
| OPUS-MT | 180 | 92.1 | 91.9 | [ , ] |
| mBART50 + m2m100 | 59 | 72.5 | 74.6 | [ , ] |
| All | 608 | 87.8 | 89.1 | [ , ] |
| Model | depth | PT | PL | DE | AR | ZH | RU | TR |
|---|---|---|---|---|---|---|---|---|
| Llama-3B | 28 | 14, 20, 24 | 14, 22, 26 | 20, 6, 22 | 16, 20, 14 | 20, 22, 12 | 14, 20, 12 | 14, 12, 6 |
| Mistral-7B | 32 | 18, 16, 12 | 18, 30, 28 | 18, 30, 28 | 10, 8, 6 | 16, 12, 14 | 18, 28, 30 | 18, 10, 26 |
| Gemma-9B | 42 | 28, 22, 26 | 26, 24, 22 | 24, 22, 28 | 22, 26, 24 | 38, 22, 36 | 24, 22, 26 | 28, 26, 24 |
| Mistral-24B | 40 | 24, 26, 28 | 36, 18, 14 | 20, 24, 34 | 24, 20, 6 | 18, 6, 8 | 26, 24, 14 | 26, 14, 12 |
| Gemma-27B | 46 | 42, 24, 30 | 22, 14, 32 | 22, 30, 18 | 24, 20, 34 | 38, 14, 36 | 24, 22, 30 | 16, 8, 32 |
| Qwen-32B | 64 | 48, 54, 12 | 51, 54, 57 | 54, 48, 45 | 51, 48, 15 | 42, 36, 57 | 51, 54, 45 | 48, 57, 54 |
| Model | fitted | scrambled pairing | beats | |
|---|---|---|---|---|
| Llama-3B | 0.821 | 0.583 | 100% | |
| Mistral-7B | 0.779 | 0.520 | 100% | |
| Gemma-9B | 0.857 | 0.730 | 100% | |
| Mistral-24B | 0.877 | 0.640 | 100% | |
| Gemma-27B | 0.839 | 0.723 | 100% | |
| Qwen-32B | 0.886 | 0.644 | 100% |
| Arm | What it tests | AUROC | vs method | success |
|---|---|---|---|---|
| baseline | nothing is applied | 0.820 | yes | |
| shift only | whether moving the cloud is enough, without turning it | 0.834 | yes | |
| rotation only | whether turning it is enough, without moving it | 0.844 | yes | |
| unconstrained linear map | whether the orthogonality constraint earns its place | 0.786 | yes | |
| scrambled pairing | whether the fact-to-fact correspondence matters, or only the two clouds | 0.621 | yes | |
| random rotation | whether any rotation would do as well as the fitted one | 0.658 | yes |