Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PII
Authors: Alexandru Nazare, Agnese Profico, Nicolò Vania, Elena Di Croce, Daria Caramanica, Davide Venditti, Elena Sofia Ruzzetti, Giancarlo A. Xompero, +1 more
Organizations: Human-Centric ART, University of Rome Tor Vergata · Department of Computer Science, University of Luxembourg · Almawave Labs, Rome
The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (TDE) attacks when prompted in non-English languages. We construct a multi-domain PII dataset comprising social media handles, email addresses, and phone numbers and translate the attack contexts into Italian, Spanish, French, and German. Our results show that TDE attacks against both English-centric and multilingual models transfer to different languages: the attacks are successful on translated prompts, even though only the original English prompt might have been included in the pre-training data. A web-presence check on a sample of the translations confirms that they are not available online. The share of English leaks recovered in other languages grows with the multilingual capability of the model, and it drops sharply when the original wording is lost, even without a change of language. This suggests that native multilingual pre-training facilitates the emergence of latent cross-linguistic bridges that simplify the retrieval of personally identifiable information (PII). We analyze the activations of multilingual large language models (LLMs) and find that different translations of the same prompt are bridged in similar representations, with the strongest alignment in the middle layers. Our results highlight a fundamental security gap in modern LLMs, necessitating more robust, language-agnostic sanitization strategies for future model alignment.
Figures & tables
PII Type
Statistic
Queries w/
Historical
Exact
Max. Fuzzy
results %
captures per query
match
score
Twitter
Mean
97.6
0.053
0
0.229
Std.
15.2
0.615
0
3.237
Min.
0
0
0
0
25th percentile
100
0
0
0
Median
100
0
0
0
TABLE I: Translated samples from Pile-CC (200 per PII type) are novel. We report descriptive statistics of the web-presence search results for Twitter handles, email addresses and phone numbers. The Queries w/ results % reports the proportion of queries returning at least one result, while Historical captures per query reports the number of historical captures identified for each query. Exact match and Max. Fuzzy score denotes the exact lexical overlap and the maximum fuzzy matching score as detailed in Section III-B1 .
PII type
Model
English
French
Italian
Spanish
German
Leak
Leak
Rel.
Generated
Leak
Rel.
Generated
Leak
Rel.
Generated
Leak
Rel.
Generated
Twitter
Llama 3.2 1B
340
255
75.0%
1788
243
71.5%
1724
260
76.5%
1893
213
62.6%
1369
Llama 3.2 3B
548
492
89.8%
2284
501
91.4%
2327
488
89.1%
2334
389
71.0%
1831
Qwen2.5 3B
468
395
84.4%
2330
372
79.5%
2360
384
82.1%
2391
298
63.7%
1693
Qwen2.5 7B
591
479
81.0%
2682
439
74.3%
2647
440
74.5%
2626
361
61.1%
2197
GPT-J 6B
1020
426
41.8%
2033
450
44.1%
2300
444
43.5%
2219
337
33.0%
1688
TABLE II: A large portion of PII are leaked also in other languages. We report the number of successfully leaked PII, together with the percentage of English leaks also observed in each language. The number of generations containing PII instances is reported for each language.
Fig. 1: Retention of leaked PII instances under a paraphrased attack, compared to strict-attack leak count, averaged across the four multilingual models (Llama 3.2 1B/3B, Qwen2.5 3B/7B) and the three English-centric models (GPT-J 6B, GPT-Neo 1.3B/2.7B)
Fig. 2: Mean cosine similarity between the representations of each strict translation and paraphrase and the English original, by layer, for each PII type.
PII type
Model
Strict
Strict
English
Target language
(FR/IT/ES)
(DE)
paraphrase
paraphrase
Twitter
Llama 3.2 1B
0.389
0.366
0.296
0.255
Llama 3.2 3B
0.375
0.367
0.279
0.239
Qwen2.5 3B
0.681
0.603
0.524
0.417
Qwen2.5 7B
0.668
0.607
0.491
0.381
Email
Llama 3.2 1B
0.340
0.338
0.309
0.255
TABLE III: Mean cosine similarity to the English representation, averaged over all layers. Strict translations are reported separately for French/Italian/Spanish (mean over the three languages) and German; paraphrases are averaged over the four target languages.
Language
Model
English
Qwen2.5-7B-Instruct
German
Qwen2.5-7B-Instruct
French
Mistral-7B-Instruct-v0.3
Italian
Minerva-7B-instruct-v1.0
Spanish
Salamandra-7B-instruct
TABLE IV: Models used for paraphrase generation.
PII type
Language
Model
BLEU
ROUGE-L
Sem. sim.
Email
English
Qwen2.5-7B-Instruct
23.51
49.21
83.43
French
Mistral-7B-Instruct-v0.3
8.21
15.02
69.23
Italian
Minerva-7B-instruct
22.08
39.94
46.99
Spanish
Salamandra-7B-instruct
6.62
18.84
41.35
German
DiscoLM_German_7b
41.89
56.41
68.05
German
Qwen2.5-7B-Instruct
47.80
67.70
90.89
TABLE V: Mean BLEU, ROUGE-L (F1), and bidirectional semantic similarity between original excerpts and their paraphrases, by PII type and language. Lower BLEU/ROUGE-L and higher semantic similarity indicate better paraphrase quality.
PII type
Model
BLEU
ROUGE-L
Sem. sim.
Email
DiscoLM_German_7b
41.89
56.41
68.05
Qwen2.5-7B-Instruct
47.80
67.70
90.89
Phone
DiscoLM_German_7b
49.31
62.21
69.85
Qwen2.5-7B-Instruct
41.98
63.94
90.70
Twitter
DiscoLM_German_7b
45.20
56.36
61.46
Qwen2.5-7B-Instruct
43.12
64.79
92.02
TABLE VI: Comparison of two candidate paraphrase models for German.
English
French
Italian
Spanish
German
Model
Strict
Paraphrase
Strict
Paraphrase
Strict
Paraphrase
Strict
Paraphrase
Strict
Paraphrase
Twitter
Llama 3.2 1B
340
68
255
56
243
16
260
36
213
25
Llama 3.2 3B
548
127
492
96
501
28
488
68
389
36
Qwen2.5 3B
468
146
395
104
372
19
384
55
298
31
Qwen2.5 7B
591
181
479
135
439
20
440
70
361
43
TABLE VII: Number of leaked PII instances for the strict translation and paraphrased attack prefixes. Results are reported for each model, language, and PII type.