Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.
Figures & tables
Figure 1: Language-budgeted unlearning. COVER selects K source languages (blue) to cover held-out languages (grey). Shaded regions group languages by script for visual orientation. Connection and spatial layout are schematic.
Axis
Size
Composition
Language–script pairs
174
18 scripts; 27 high-, 46 medium-, 101 low-resource; scored in both roles
Question forms
25
ETPC atomic paraphrase types in seven meta-categories ( Kovatchev et al., 2018 ) ; seven types translated into 40 languages (1,690 rows per language)
Fine-tuned checkpoints
103
81 monolingual (27 languages per model family) and 22 multilingual configurations over Tiny Aya ( Salamanca et al., 2026 ) , Qwen3 ( Yang et al., 2025 ) and Llama 3 ( Grattafiori et al., 2024 ) , each with a matched retain-only oracle
Unlearned checkpoints
287
243 monolingual endpoints ( 81×3 methods: SimNPO ( Fan et al., 2025 ) , RMU ( Li et al., 2024 ) , GradDiff ( Dorna et al., 2025 ) ) and 44 Aya trilingual endpoints, stopped at matched source forgetting
Table 1: The Cross-lingual Unlearning Tensor at a glance. Coverage across language, expression, and model axes; scored routes are defined in the text.
Figure 2: Uneven RMU transfer across language routes. Points are evaluation languages; acquired access is A=pFT−pO and removal is U=pFT−pU . Shading marks U<A . (a) Means over monolingual checkpoints. (b) Qwen English, Qwen Russian, and Aya Arabic fine-tunes on ℓ→a .
(A) Acquisition: likelihood gain over base
Model
a→a
ℓ→a
ℓ→ℓ
ℓ→ℓ with sufficient access
Aya
+0.85
+0.49
+0.14
65.3%
Qwen
+0.84
+0.47
+0.10
65.5%
Llama
+0.76
+0.45
+0.03
29.6%
Mean
+0.82
+0.47
+0.09
53.5%
(B) Unlearning: oracle-normalized removal (%)
Table 2: Acquisition and unlearning transfer across language routes. (A) Gains over the base model after fine-tuning in a ; sufficient access requires pFT≥0.10 and gain ≥0.05 . (B) Oracle-normalized removal after unlearning in a ; final columns report languages reaching ≥90% removal and worst-language residual. Values above 100 indicate suppression below the retain-only oracle.
Residual access 100R↓
Gain over uniform ↑
Ranking consistency
Inventory
Model
Best
Uniform
Worst
Best (pts)
Worst (pts)
Best (%)
ρ
Script-diverse
Qwen
33.59
41.94
52.93
+8.35
−10.99
19.9
0.813
Aya
31.38
41.38
54.17
+10.00
−12.79
24.2
0.850
Llama
46.15
53.69
62.26
+7.53
−8.57
14.0
0.663
Low-resource
Qwen
34.32
48.60
65.92
+14.28
−17.32
29.4
0.886
Aya
28.89
39.85
58.97
+10.96
−19.12
27.5
0.875
Table 3: Source choice affects forgetting, with inventory-dependent ranking stability. Residuals average five training orders; best and worst triples minimize and maximize this mean. Uniform averages all 20 triples. Gains are point or relative reductions from uniform. ρ is the median pairwise Spearman correlation across order-specific rankings.
(a) TOFU, low-resource inventory, two held-out requests
Reduction in held-out 100R↑
Model / request
Uniform 100R↓
Fixed coverage
Historical best
COVER
Aya / A
39.23
+10.70± 1.62
+7.03± 2.43
= Fixed
Aya / B
36.57
+9.84± 0.89
+9.36± 2.10
= Fixed
Qwen / A
49.94
+6.72± 3.13
+5.65± 1.87
= Fixed
Qwen / B
47.06
+6.14± 4.29
+4.38± 2.96
= Fixed
Table 4: Source selection across deletion requests and datasets. Gains are point reductions in 100R from uniform over all 20 triples; extra retain damage is in retained-probability points. (a) TOFU: mean ± SD over five orders; “= Fixed” marks identical subsets. (b) LUME request B: two realizations; all-six includes sources. (c) LORELEI: means over two orders; Δ chrF++ is on retained generations. Historical best reuses the best subset from an earlier request.
Model
Policy
ES ↓
Full leakage ↓
Any disclosure ↓
Retained recovery ↑
Gibberish ↓
Code- switching ↓
Aya
Controls
0.2572
29.35
52.28
69.48
1.52
6.70
Aya
Individual top-3
–
29.76
52.15
70.33
4.06
8.76
Aya
COVER
0.2319
25.57
47.13
71.44
4.39
9.41
Qwen
Controls
0.2433
27.32
50.93
75.24
2.41
7.92
Qwen
Individual top-3
–
25.22
49.63
72.93
1.72
8.22
Qwen
COVER
0.2215
25.22
48.13
71.60
2.91
6.52
Table 5: Generated-answer evaluation on the script-diverse inventory. ES and behavior (%), averaged over three orders and six native routes (retention: Aya six, Qwen/Llama five excluding Arabic). Random controls: EN/AR/ZH, EN/JA/ES, AR/JA/RU. Individual top-3 selects the strongest singleton sources. Metric definitions: subsection A.6 .
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Script
Resource tier
Count
Languages
Latin [-1pt] Latn ⋅ 121
High
19
cat ces dan deu en est fin fra hrv hun ita nld nob pol por slk spa swe vie
Medium
25
afr als azj bos ceb cym epo eus gle glg hau ind isl jav lit lvs mlt nno ron slv swh tgl tur uzn zsm
Low
77
ace aka arb ast bam ban bem bjn bug cjk crh dik dyu ewe fao fon fur fuv gaz gla grn hat ibo ilo kab kac kam kbp kea kin kmb knc kon lij lim lmo ltg ltz lua lug luo lus min mos mri nso nus nya oci pag pap plt quy run sag scn smo sna som sot srd ssw sun szl taq tpi tsn tso tuk tum twi umb vec war wol xho zul
Arabic [-1pt] Arab ⋅ 20
High
1
arb
Medium
3
arz pes urd
Low
16
ace acm acq aeb ajp apc ars ary azb bjn kas knc min pbt prs snd
Appendix
Table 6: Language coverage and paraphrases. The benchmark contains 174 language–script pairs spanning 18 scripts. The 40 languages shown in green have translated paraphrase variants available.
Category
Paraphrase type
Definition
Morphology
Inflectional changes
Changes grammatical inflection, such as number, tense, or person.
Modal verb changes
Expresses the same proposition using a different modal construction.
Derivational changes
Uses a morphologically related word.
Example How has Elvin Mammadov contributed to fiction literature? How has Elvin Mammadov contributed to literary fiction?
Lexicon
Spelling changes
Uses alternative spellings of the same word.
Same polarity substitution (habitual)
Replaces an expression with an applicable synonym.
Appendix
Table 7: The 25 atomic paraphrase types in the ETPC ( Kovatchev et al., 2018 ) . Examples are given per category; the second line is a paraphrase of the first.
Checkpoint inventory
Model family
Training regime
FT
Retain-90 oracle
Unlearned checkpoints
Tiny Aya
Monolingual
27
27
81
Trilingual endpoint cohort
7
7
44
Additional five-/six-language mixtures
5
5
—
Qwen3
Monolingual
27
27
81
Additional five-/six-language mixtures
5
5
—
Appendix
Table 8: Model and checkpoint coverage. Panel A summarizes the saved checkpoints by model family. Panel B gives the multilingual language sets.
Inventory
Acquisition parent
Mean full recovery
Lowest full recovery
Mean target-only
Original / script-diverse
Tiny Aya Global
88.56
76.00 (JA)
99.39
Qwen3-4B-Instruct-2507
91.17
76.33 (JA)
99.39
Llama-3.2-3B-Instruct
93.61
90.67 (AR)
99.33
Latin-script
Tiny Aya Global
92.33
86.00 (TR)
99.67
Qwen3-4B-Instruct-2507
94.33
92.00 (DE, ID)
98.67
Llama-3.2-3B-Instruct
97.00
92.00 (SW)
99.67
Appendix
Table 9: Capability of the language-set specific model families before unlearning. All entries are percentages.
Language
Tiny Aya Global
Qwen3 4B
Llama 3.2 3B
Amharic
48.6
36.1
26.8
Arabic
63.0
85.4
58.3
Bengali
59.2
72.0
44.0
Czech
68.7
84.3
55.9
German
67.8
88.7
51.4
Greek
68.2
82.6
43.4
Appendix
Table 10: Released checkpoint Belebele accuracy across the 46 evaluated languages.
Label
Agree (%)
Cohen’s κ
Rate difference (pts)
Fact status ( none / partial / full )
81.8
0.64 [0.50, 0.77]
–
Full leakage
94.7
0.86 [0.69, 0.98]
−4.1 [ −9.9 , +0.4 ]
Any disclosure
88.6
0.77 [0.64, 0.87]
−11.4 [ −17.9 , −6.2 ]
Appendix
Table 11: Agreement between Qwen3.5-27B and a blind human rater on 192 facts. Brackets: 95% stratified bootstrap intervals. Rate difference is the judge’s rate minus the human’s.
Translation model
Languages judged
Mean win rate vs. NLLB (%) ↑
GPT-5.2-chat
40
94.8%
GPT-5
30
94.4%
DeepSeek-V3
35
88.9%
GPT-4.1
40
88.3%
GPT-5.1
40
88.0%
GPT-5-mini
35
87.2%
Appendix
Table 12: For each language and translation model, win rate is the fraction of items where the model translation won relative to NLLB. GPT-5.2-chat had the highest reported mean while also providing full language coverage.
Language
Mean.
Answer cons.
Flu.
Names/ num.
zho_Hant
3.44
3.25
3.76
2.17
zho_Hans
2.98
2.67
3.52
1.84
ukr_Cyrl
2.45
1.96
3.21
2.12
heb_Hebr
2.58
2.20
3.20
1.76
arb_Arab
2.27
1.82
3.03
1.83
jpn_Jpan
2.38
1.85
3.31
1.37
Appendix
Table 13: GPT-5.2-chat translations beat NLLB on every criterion in every judged language. Judge grades (1–5, against the English gold) for meaning, answer consistency, fluency and names/numbers; each cell is the GPT-5.2-chat grade minus the NLLB grade, so positive favors GPT-5.2-chat. Languages are sorted by the mean of the four differences.
Figure 5: Reference-free translation quality. Each point is one language’s mean COMETKiwi score over the sampled QA pairs, a pair score averaging its question and answer, sorted from lowest to highest. The panel mean is 0.709 and the median 0.768
Figure 6: Mean paired COMETKiwi difference (our translation minus NLLB), in COMETKiwi score units, over 200 aligned segments per language, for the ten languages with the largest advantage and the ten with the smallest. Positive values favor our translation; negative values favor NLLB.
Figure 7: NLLB round-trip fidelity over 54 languages. Distributions over QA pairs of chrF++ between the original English and the NLLB back-translation (left, median 69.5) and of embedding cosine similarity between the same two texts (right, median 0.955). Dotted lines mark the median and dashed lines the outlier threshold of the screen in Appendix B.4 .
Metric
Screening rule
Languages flagged
Direct COMETKiwi
Score <0.300964
0
GPT-5.2 − NLLB COMETKiwi
Score <−0.062456
1
Back-translation chrF++
Score <42.834452
4
Back-translation semantic similarity
Score <0.789314
7
Back-translation pair-failure rate
Rate >50%
19
Appendix
Table 14: Screening rules and the number of languages flagged by each. Cutoffs for the four primary metrics are set at the median minus 3×1.4826×MAD ; the pair-failure screen uses the fixed criterion that more than half of QA pairs fail.
Classification
Languages
Language names
Outlier on two or more metrics
4
Chokwe, Dyula, Fon, Tosk Albanian
Outlier on one metric
3
Central Kanuri (Arabic), Jingpho, Southwestern Dinka
NLLB diagnostics unavailable
2
Minangkabau (Arabic), Modern Standard Arabic (Romanized)
No issues
152
See supplementary table
Appendix
Table 15: Pruned languages are almost exclusively low resource but only a handful were ultimately ruled out. Outlier metrics are the four direct and back-translation metrics in Table 14 .
(a) Source languages: gain over 173 other language–script pairs
Aya
Qwen
Llama
Rank
Language
Gain
Language
Gain
Language
Gain
1
Spanish
0.1664
Spanish
0.1335
Spanish
0.0667
2
Norwegian Bokmål
0.1621
Vietnamese
0.1264
German
0.0625
3
Indonesian
0.1616
Norwegian Bokmål
0.1262
Vietnamese
0.0596
4
Vietnamese
0.1558
Indonesian
0.1259
Thai
0.0531
Appendix
Table 16: Acquisition transfer rankings. Mean fine-tuned-minus-base gold-answer likelihood gain per model family (common 21-language acquisition cohort, 50 forget facts per cell; no oracle subtraction or eligibility filter). (a) Top ten acquisition (source) languages by mean gain over the 173 other language–script pairs, with question and answer in the target language. (b) Top ten receiving question languages, with answers kept in the acquisition language. (c) Top ten receiving languages, with both question and answer in the named language.
Method
Most frequent bottleneck languages: checkpoints /70 (Aya/Qwen/Llama)
North Levantine Arabic 4 (1/3/0), Russian 4 (1/1/2), Simplified Chinese 4 (0/3/1), Wolof 3 (3/0/0), Ta’izzi-Adeni Arabic 2 (1/0/1), Moroccan Arabic 2 (0/0/2), Crimean Tatar 2 (2/0/0), English 2 (0/0/2), Galician 2 (0/1/1), Occitan 2 (1/1/0)
12 (2/1/9)
RMU
Russian 27 (8/4/15), English 10 (0/2/8), Arabic 6 (6/0/0), Ta’izzi-Adeni Arabic 4 (2/2/0), Simplified Chinese 4 (0/4/0)
0
Appendix
Table 17: A few languages are the recurring bottlenecks of unlearning transfer. For each acquisition-language unlearned checkpoint, the other language with the largest positive remaining acquired-access fraction R=(pU−pO)/(pFT−pO) , with question and answer both in that language, is its bottleneck. Entries count how often each language is the bottleneck over the 70 checkpoints (23 Aya, 21 Qwen, 26 Llama; per-model counts in parentheses), top five per method with ties. “None” counts checkpoints in which every eligible alternative is at or below the oracle.
Unlearning method
Aya
Qwen
Llama
Question translated; answer stays in the training language
SimNPO
0.99
0.98
0.99
GradDiff
0.99
0.97
0.98
RMU
0.96
0.86
0.92
Question and answer translated
SimNPO
0.86
0.76
0.71
Appendix
Table 18: Languages that gain the most access during learning also lose the most during unlearning. Median within-checkpoint Spearman rank correlation between fine-tuning gain over base pFT−pB and the subsequent correct-answer likelihood decrease pFT−pU after unlearning in the same language, across 21 training-language checkpoints per model and the two answer-language settings shown.
Mean win rate (%)
Rank after SimNPO
Rank
Paraphrase type
QA
Before
After
Aya
Qwen
Llama
1
Modal verb
28
77.7
67.6
1
1
1
2
Subordination
63
54.5
58.8
3
3
2
3
Sentence modality
100
53.7
58.1
2
2
3
4
Same-polarity habitual
91
52.1
47.4
4
5
6
5
Same-polarity contextual
92
53.0
45.7
8
7
4
Appendix
Table 19: Modal-verb, subordination and sentence-modality paraphrases leave the most access after unlearning. English paraphrase types ranked by matched-question likelihood win rate, averaged over competing types and model families. Before/After denote acquisition/post-SimNPO; family ranks refer to After (1 = highest post-unlearning likelihood win rate).
Method
Matched-form advantage (NLL) ↑
95% interval
SimNPO
6.05
[5.83, 6.27]
RMU
8.26
[7.66, 8.80]
Appendix
Table 20: Question form changes which access routes are suppressed. The matched-form advantage is the mean additional answer-NLL increase when training and evaluation both use native Hindi questions or both use Romanized Hindi, compared with crossing the two forms. The contrast is in answer-NLL units; positive values indicate greater suppression with matched forms. Answers remain in native Hindi. The 16 runs cross two model families, two seeds, two methods and two training forms at a fixed 100-update budget. Intervals bootstrap the 20 authors in the 318-fact common assessment panel.
Forget generations
Forget supervision
Full leakage ↓
Gibberish ↓
Retain damage ↓
English only ( K=1 )
0.553
0.031
0.129
Best triple ( K=3 )
0.196
0.129
0.105
All six ( K=6 )
0.137
0.169
0.130
Appendix
Table 21: All-language forget supervision removes best but degrades every language’s generations. Forget supervision in K of the six acquired languages, averaged over the nine development parents. Supervising one language leaves the fact recoverable in the others; supervising all six removes it, with more than five times the degenerate output.
Ranking agreement
Leave-one-order-out
Model / request
Cross-request ρ
Within-request ρ
Shared top five
Four-to-one gain mean [min, max]
Positive held-out orders
Qwen / A
0.929
0.802
3/5
+7.11 [ +0.51 , +11.97 ]
5/5
Qwen / B
0.668
0.760
1/5
+5.72 [ +2.69 , +10.75 ]
5/5
Tiny Aya / A
0.923
0.840
4/5
+10.55 [ +5.77 , +13.36 ]
5/5
Tiny Aya / B
0.950
0.838
5/5
+9.84 [ +7.03 , +12.88 ]
5/5
Llama / A
0.913
0.776
3/5
+9.24 [ +4.16 , +13.72 ]
5/5
Appendix
Table 23: Source-set rankings recur across training orders and across disjoint forget requests (TOFU, low-resource inventory, K=3 ). Cross-request ρ is the Spearman correlation between the rankings of two disjoint forget requests after averaging each subset over its five orders; within-request ρ is the median of the ten pairwise correlations between order-specific rankings; shared top five counts subsets in both requests’ top five. Four-to-one selects the lowest-residual subset on four orders of the same request and evaluates it on the held-out fifth (gain in 100R over uniform).
Allocation
Full ↓
Any ↓
Retain ↑
Gibberish ↓
Target-only ↑
Qwen: EN/AR/JA
Uniform
34.33
57.02
82.78
2.26
87.94
Calibrated
29.61
52.07
83.78
3.91
90.31
English-heavy
31.67
53.94
80.67
3.11
88.24
AR/JA-permuted
30.94
52.74
85.00
2.17
86.72
Llama: EN/AR/JA
Appendix
Table 24: Source-weighting comparisons. Rates (%) averaged over three orders on TOFU. Retain uses all six languages in (a) and the indicated held-out languages in (b)
Figure 8: Good and bad source sets recur across training orders. Each panel ranks the 20 three-language subsets (rows, ordered by their mean rank) by held-out residual access under each of five training orders (columns); 1 is best and 20 worst. The script-diverse inventory is English, Arabic, Japanese, Chinese, Russian and Spanish; the low-resource inventory is English, Arabic, Hausa, Hindi, Nepali and Swahili. The top and bottom of each ranking recur across orders while the middle is less stable, and the low-resource inventory (b, d) is the more stable of the two, in line with the consistency values of Table 3 .
Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. However, unlearning research remains heavily English-centric. We study multilingual unlearning by extending the TOFU benchmark to five languages, and fine-tune, unlearn, and query our models with different permutations of languages. We find that unlearning transfer, the ability of an unlearned model to "forget" facts in languages other than the unlearning language, is highly variable: e.g., it is strongest between languages sharing scripts and families, and we show that the unlearning language predicts which query languages are most likely to yield the strongest transfer. Layer-wise analysis reveals that unlearning leaves the shared cross-lingual latent space largely intact in early layers, instead operating primarily in later decoding layers. This suggests that unlearning does not truly erase knowledge, but rather induces superficial suppression. Exploiting this structure, a single inference-time steering direction reverses much of this suppression across languages, recovering 50% (Qwen) and 90% (Gemma) of the unlearned knowledge.
Chaoyi Xiang, Olga Ohrimenko, Benjamin I. P. Rubinstein +1
School of Computing and Information Systems, The University of Melbourne, Melbourne, Australia.
While LLMs are increasingly used in commercial services, they pose privacy risks such as leakage of sensitive personally identifiable information (PII). For LLMs trained on multilingual corpora, Multilingual Machine Unlearning (MMU) aims to remove information across multiple languages. However, prior MMU evaluations fail to capture such cross-linguistic distribution of information, being largely limited to direct extensions of per-language evaluation protocols. To this end, we propose two metrics to evaluate the information spread across languages: the Knowledge Separability Score (KSS) and the Knowledge Persistence Score (KPS). KSS measures the overall unlearning quality across multiple languages, while KPS more specifically aims to assess consistent removal of information among different language pairs. We evaluated various unlearning methods in the multilingual setting with these metrics and conducted comprehensive analyses. Through our investigation, we provide insights into unique phenomena exclusive to MMU and offer a new perspective on MMU evaluation.
Kyomin Hwang, Hyeonjin Kim, Sangyeon Cho +1
GSCST, Seoul National University · Department of Artificial Intelligence, Chung-Ang University · Korean Surgical Researcher Foundation, Republic of Korea +1
Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks include mainly single-hop questions and a narrow set of multi-hop questions. Although effective, they still face two challenges. (1) Knowledge is not isolated, whereby diverse multi-hop reasoning paths can potentially induce knowledge leakage than normal queries. (2) Unlearning may be fragile: unlearned knowledge can be partially recovered through recovery attacks such as lightweight post-unlearning adaptation, making static evaluation insufficient. Therefore, in this paper, we introduce \unlearning as a novel benchmark to understand robust LLM knowledge removal across diverse reasoning paths and recovery attacks. We experiment with this benchmark on 3 models, 6 unlearning methods, and 2 carefully curated datasets. Results show that existing methods are vulnerable to multi-hop reasoning paths and recovery attacks. We further explore the trade-off among forget quality, robustness, and model utility for LLM unlearning.