Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
Organizations: Johns Hopkins University
Abstract
Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.
Figures & tables
| Axis | Size | Composition |
| Language–script pairs | 174 | 18 scripts; 27 high-, 46 medium-, 101 low-resource; scored in both roles |
| Question forms | 25 | ETPC atomic paraphrase types in seven meta-categories ( Kovatchev et al., 2018 ) ; seven types translated into 40 languages (1,690 rows per language) |
| Fine-tuned checkpoints | 103 | 81 monolingual (27 languages per model family) and 22 multilingual configurations over Tiny Aya ( Salamanca et al., 2026 ) , Qwen3 ( Yang et al., 2025 ) and Llama 3 ( Grattafiori et al., 2024 ) , each with a matched retain-only oracle |
| Unlearned checkpoints | 287 | 243 monolingual endpoints ( methods: SimNPO ( Fan et al., 2025 ) , RMU ( Li et al., 2024 ) , GradDiff ( Dorna et al., 2025 ) ) and 44 Aya trilingual endpoints, stopped at matched source forgetting |
| (A) Acquisition: likelihood gain over base | |||||
| Model | with sufficient access | ||||
| Aya | 65.3% | ||||
| Qwen | 65.5% | ||||
| Llama | 29.6% | ||||
| Mean | 53.5% | ||||
| (B) Unlearning: oracle-normalized removal (%) | |||||
| Residual access | Gain over uniform | Ranking consistency | ||||||
| Inventory | Model | Best | Uniform | Worst | Best (pts) | Worst (pts) | Best (%) | |
| Script-diverse | Qwen | 33.59 | 41.94 | 52.93 | 19.9 | 0.813 | ||
| Aya | 31.38 | 41.38 | 54.17 | 24.2 | 0.850 | |||
| Llama | 46.15 | 53.69 | 62.26 | 14.0 | 0.663 | |||
| Low-resource | Qwen | 34.32 | 48.60 | 65.92 | 29.4 | 0.886 | ||
| Aya | 28.89 | 39.85 | 58.97 | 27.5 | 0.875 | |||
| (a) TOFU, low-resource inventory, two held-out requests | ||||
| Reduction in held-out | ||||
| Model / request | Uniform | Fixed coverage | Historical best | COVER |
| Aya / A | 39.23 | 1.62 | 2.43 | = Fixed |
| Aya / B | 36.57 | 0.89 | 2.10 | = Fixed |
| Qwen / A | 49.94 | 3.13 | 1.87 | = Fixed |
| Qwen / B | 47.06 | 4.29 | 2.96 | = Fixed |
| Model | Policy | ES | Full leakage | Any disclosure | Retained recovery | Gibberish | Code- switching |
| Aya | Controls | 0.2572 | 29.35 | 52.28 | 69.48 | 1.52 | 6.70 |
| Aya | Individual top-3 | – | 29.76 | 52.15 | 70.33 | 4.06 | 8.76 |
| Aya | COVER | 0.2319 | 25.57 | 47.13 | 71.44 | 4.39 | 9.41 |
| Qwen | Controls | 0.2433 | 27.32 | 50.93 | 75.24 | 2.41 | 7.92 |
| Qwen | Individual top-3 | – | 25.22 | 49.63 | 72.93 | 1.72 | 8.22 |
| Qwen | COVER | 0.2215 | 25.22 | 48.13 | 71.60 | 2.91 | 6.52 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Script | Resource tier | Count | Languages |
| Latin [-1pt] Latn 121 | High | 19 | cat ces dan deu en est fin fra hrv hun ita nld nob pol por slk spa swe vie |
| Medium | 25 | afr als azj bos ceb cym epo eus gle glg hau ind isl jav lit lvs mlt nno ron slv swh tgl tur uzn zsm | |
| Low | 77 | ace aka arb ast bam ban bem bjn bug cjk crh dik dyu ewe fao fon fur fuv gaz gla grn hat ibo ilo kab kac kam kbp kea kin kmb knc kon lij lim lmo ltg ltz lua lug luo lus min mos mri nso nus nya oci pag pap plt quy run sag scn smo sna som sot srd ssw sun szl taq tpi tsn tso tuk tum twi umb vec war wol xho zul | |
| Arabic [-1pt] Arab 20 | High | 1 | arb |
| Medium | 3 | arz pes urd | |
| Low | 16 | ace acm acq aeb ajp apc ars ary azb bjn kas knc min pbt prs snd |
| Category | Paraphrase type | Definition |
| Morphology | Inflectional changes | Changes grammatical inflection, such as number, tense, or person. |
| Modal verb changes | Expresses the same proposition using a different modal construction. | |
| Derivational changes | Uses a morphologically related word. | |
| Example How has Elvin Mammadov contributed to fiction literature? How has Elvin Mammadov contributed to literary fiction? | ||
| Lexicon | Spelling changes | Uses alternative spellings of the same word. |
| Same polarity substitution (habitual) | Replaces an expression with an applicable synonym. | |
| Checkpoint inventory | ||||
| Model family | Training regime | FT | Retain-90 oracle | Unlearned checkpoints |
| Tiny Aya | Monolingual | 27 | 27 | 81 |
| Trilingual endpoint cohort | 7 | 7 | 44 | |
| Additional five-/six-language mixtures | 5 | 5 | — | |
| Qwen3 | Monolingual | 27 | 27 | 81 |
| Additional five-/six-language mixtures | 5 | 5 | — | |
| Inventory | Acquisition parent | Mean full recovery | Lowest full recovery | Mean target-only |
| Original / script-diverse | Tiny Aya Global | 88.56 | 76.00 (JA) | 99.39 |
| Qwen3-4B-Instruct-2507 | 91.17 | 76.33 (JA) | 99.39 | |
| Llama-3.2-3B-Instruct | 93.61 | 90.67 (AR) | 99.33 | |
| Latin-script | Tiny Aya Global | 92.33 | 86.00 (TR) | 99.67 |
| Qwen3-4B-Instruct-2507 | 94.33 | 92.00 (DE, ID) | 98.67 | |
| Llama-3.2-3B-Instruct | 97.00 | 92.00 (SW) | 99.67 |
| Language | Tiny Aya Global | Qwen3 4B | Llama 3.2 3B |
| Amharic | 48.6 | 36.1 | 26.8 |
| Arabic | 63.0 | 85.4 | 58.3 |
| Bengali | 59.2 | 72.0 | 44.0 |
| Czech | 68.7 | 84.3 | 55.9 |
| German | 67.8 | 88.7 | 51.4 |
| Greek | 68.2 | 82.6 | 43.4 |
| Label | Agree (%) | Cohen’s | Rate difference (pts) |
| Fact status ( none / partial / full ) | 81.8 | 0.64 [0.50, 0.77] | – |
| Full leakage | 94.7 | 0.86 [0.69, 0.98] | [ , ] |
| Any disclosure | 88.6 | 0.77 [0.64, 0.87] | [ , ] |
| Translation model | Languages judged | Mean win rate vs. NLLB (%) |
| GPT-5.2-chat | 40 | 94.8% |
| GPT-5 | 30 | 94.4% |
| DeepSeek-V3 | 35 | 88.9% |
| GPT-4.1 | 40 | 88.3% |
| GPT-5.1 | 40 | 88.0% |
| GPT-5-mini | 35 | 87.2% |
| Language | Mean. | Answer cons. | Flu. | Names/ num. |
| zho_Hant | 3.44 | 3.25 | 3.76 | 2.17 |
| zho_Hans | 2.98 | 2.67 | 3.52 | 1.84 |
| ukr_Cyrl | 2.45 | 1.96 | 3.21 | 2.12 |
| heb_Hebr | 2.58 | 2.20 | 3.20 | 1.76 |
| arb_Arab | 2.27 | 1.82 | 3.03 | 1.83 |
| jpn_Jpan | 2.38 | 1.85 | 3.31 | 1.37 |
| Metric | Screening rule | Languages flagged |
| Direct COMETKiwi | Score | 0 |
| GPT-5.2 NLLB COMETKiwi | Score | 1 |
| Back-translation chrF++ | Score | 4 |
| Back-translation semantic similarity | Score | 7 |
| Back-translation pair-failure rate | Rate | 19 |
| Classification | Languages | Language names |
| Outlier on two or more metrics | 4 | Chokwe, Dyula, Fon, Tosk Albanian |
| Outlier on one metric | 3 | Central Kanuri (Arabic), Jingpho, Southwestern Dinka |
| NLLB diagnostics unavailable | 2 | Minangkabau (Arabic), Modern Standard Arabic (Romanized) |
| No issues | 152 | See supplementary table |
| (a) Source languages: gain over 173 other language–script pairs | ||||||
| Aya | Qwen | Llama | ||||
| Rank | Language | Gain | Language | Gain | Language | Gain |
| 1 | Spanish | 0.1664 | Spanish | 0.1335 | Spanish | 0.0667 |
| 2 | Norwegian Bokmål | 0.1621 | Vietnamese | 0.1264 | German | 0.0625 |
| 3 | Indonesian | 0.1616 | Norwegian Bokmål | 0.1262 | Vietnamese | 0.0596 |
| 4 | Vietnamese | 0.1558 | Indonesian | 0.1259 | Thai | 0.0531 |
| Method | Most frequent bottleneck languages: checkpoints /70 (Aya/Qwen/Llama) | None |
| SimNPO | Kannada 5 (0/0/5), North Levantine Arabic 4 (1/2/1), Kamba 4 (2/0/2), Galician 3 (0/1/2), Gujarati 3 (3/0/0), Russian 3 (2/1/0), Simplified Chinese 3 (0/2/1) | 0 |
| GradDiff | North Levantine Arabic 4 (1/3/0), Russian 4 (1/1/2), Simplified Chinese 4 (0/3/1), Wolof 3 (3/0/0), Ta’izzi-Adeni Arabic 2 (1/0/1), Moroccan Arabic 2 (0/0/2), Crimean Tatar 2 (2/0/0), English 2 (0/0/2), Galician 2 (0/1/1), Occitan 2 (1/1/0) | 12 (2/1/9) |
| RMU | Russian 27 (8/4/15), English 10 (0/2/8), Arabic 6 (6/0/0), Ta’izzi-Adeni Arabic 4 (2/2/0), Simplified Chinese 4 (0/4/0) | 0 |
| Unlearning method | Aya | Qwen | Llama |
| Question translated; answer stays in the training language | |||
| SimNPO | 0.99 | 0.98 | 0.99 |
| GradDiff | 0.99 | 0.97 | 0.98 |
| RMU | 0.96 | 0.86 | 0.92 |
| Question and answer translated | |||
| SimNPO | 0.86 | 0.76 | 0.71 |
| Mean win rate (%) | Rank after SimNPO | ||||||
| Rank | Paraphrase type | QA | Before | After | Aya | Qwen | Llama |
| 1 | Modal verb | 28 | 77.7 | 67.6 | 1 | 1 | 1 |
| 2 | Subordination | 63 | 54.5 | 58.8 | 3 | 3 | 2 |
| 3 | Sentence modality | 100 | 53.7 | 58.1 | 2 | 2 | 3 |
| 4 | Same-polarity habitual | 91 | 52.1 | 47.4 | 4 | 5 | 6 |
| 5 | Same-polarity contextual | 92 | 53.0 | 45.7 | 8 | 7 | 4 |
| Method | Matched-form advantage (NLL) | 95% interval |
| SimNPO | 6.05 | [5.83, 6.27] |
| RMU | 8.26 | [7.66, 8.80] |
| Forget generations | |||
| Forget supervision | Full leakage | Gibberish | Retain damage |
| English only ( ) | 0.553 | 0.031 | 0.129 |
| Best triple ( ) | 0.196 | 0.129 | 0.105 |
| All six ( ) | 0.137 | 0.169 | 0.130 |
| Ranking agreement | Leave-one-order-out | ||||
| Model / request | Cross-request | Within-request | Shared top five | Four-to-one gain mean [min, max] | Positive held-out orders |
| Qwen / A | 0.929 | 0.802 | 3/5 | [ , ] | 5/5 |
| Qwen / B | 0.668 | 0.760 | 1/5 | [ , ] | 5/5 |
| Tiny Aya / A | 0.923 | 0.840 | 4/5 | [ , ] | 5/5 |
| Tiny Aya / B | 0.950 | 0.838 | 5/5 | [ , ] | 5/5 |
| Llama / A | 0.913 | 0.776 | 3/5 | [ , ] | 5/5 |
| Allocation | Full | Any | Retain | Gibberish | Target-only |
| Qwen: EN/AR/JA | |||||
| Uniform | 34.33 | 57.02 | 82.78 | 2.26 | 87.94 |
| Calibrated | 29.61 | 52.07 | 83.78 | 3.91 | 90.31 |
| English-heavy | 31.67 | 53.94 | 80.67 | 3.11 | 88.24 |
| AR/JA-permuted | 30.94 | 52.74 | 85.00 | 2.17 | 86.72 |
| Llama: EN/AR/JA | |||||