Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PII
Authors: Alexandru Nazare, Agnese Profico, Nicolò Vania, Elena Di Croce, Daria Caramanica, Davide Venditti, Elena Sofia Ruzzetti, Giancarlo A. Xompero, +1 more
Organizations: Human-Centric ART, University of Rome Tor Vergata · Department of Computer Science, University of Luxembourg · Almawave Labs, Rome
The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (TDE) attacks when prompted in non-English languages. We construct a multi-domain PII dataset comprising social media handles, email addresses, and phone numbers and translate the attack contexts into Italian, Spanish, French, and German. Our results show that TDE attacks against both English-centric and multilingual models transfer to different languages: the attacks are successful on translated prompts, even though only the original English prompt might have been included in the pre-training data. A web-presence check on a sample of the translations confirms that they are not available online. The share of English leaks recovered in other languages grows with the multilingual capability of the model, and it drops sharply when the original wording is lost, even without a change of language. This suggests that native multilingual pre-training facilitates the emergence of latent cross-linguistic bridges that simplify the retrieval of personally identifiable information (PII). We analyze the activations of multilingual large language models (LLMs) and find that different translations of the same prompt are bridged in similar representations, with the strongest alignment in the middle layers. Our results highlight a fundamental security gap in modern LLMs, necessitating more robust, language-agnostic sanitization strategies for future model alignment.
Figures & tables
PII Type
Statistic
Queries w/
Historical
Exact
Max. Fuzzy
results %
captures per query
match
score
Twitter
Mean
97.6
0.053
0
0.229
Std.
15.2
0.615
0
3.237
Min.
0
0
0
0
25th percentile
100
0
0
0
Median
100
0
0
0
TABLE I: Translated samples from Pile-CC (200 per PII type) are novel. We report descriptive statistics of the web-presence search results for Twitter handles, email addresses and phone numbers. The Queries w/ results % reports the proportion of queries returning at least one result, while Historical captures per query reports the number of historical captures identified for each query. Exact match and Max. Fuzzy score denotes the exact lexical overlap and the maximum fuzzy matching score as detailed in Section III-B1 .
PII type
Model
English
French
Italian
Spanish
German
Leak
Leak
Rel.
Generated
Leak
Rel.
Generated
Leak
Rel.
Generated
Leak
Rel.
Generated
Twitter
Llama 3.2 1B
340
255
75.0%
1788
243
71.5%
1724
260
76.5%
1893
213
62.6%
1369
Llama 3.2 3B
548
492
89.8%
2284
501
91.4%
2327
488
89.1%
2334
389
71.0%
1831
Qwen2.5 3B
468
395
84.4%
2330
372
79.5%
2360
384
82.1%
2391
298
63.7%
1693
Qwen2.5 7B
591
479
81.0%
2682
439
74.3%
2647
440
74.5%
2626
361
61.1%
2197
GPT-J 6B
1020
426
41.8%
2033
450
44.1%
2300
444
43.5%
2219
337
33.0%
1688
TABLE II: A large portion of PII are leaked also in other languages. We report the number of successfully leaked PII, together with the percentage of English leaks also observed in each language. The number of generations containing PII instances is reported for each language.
Fig. 1: Retention of leaked PII instances under a paraphrased attack, compared to strict-attack leak count, averaged across the four multilingual models (Llama 3.2 1B/3B, Qwen2.5 3B/7B) and the three English-centric models (GPT-J 6B, GPT-Neo 1.3B/2.7B)
Fig. 2: Mean cosine similarity between the representations of each strict translation and paraphrase and the English original, by layer, for each PII type.
PII type
Model
Strict
Strict
English
Target language
(FR/IT/ES)
(DE)
paraphrase
paraphrase
Twitter
Llama 3.2 1B
0.389
0.366
0.296
0.255
Llama 3.2 3B
0.375
0.367
0.279
0.239
Qwen2.5 3B
0.681
0.603
0.524
0.417
Qwen2.5 7B
0.668
0.607
0.491
0.381
Email
Llama 3.2 1B
0.340
0.338
0.309
0.255
TABLE III: Mean cosine similarity to the English representation, averaged over all layers. Strict translations are reported separately for French/Italian/Spanish (mean over the three languages) and German; paraphrases are averaged over the four target languages.
Language
Model
English
Qwen2.5-7B-Instruct
German
Qwen2.5-7B-Instruct
French
Mistral-7B-Instruct-v0.3
Italian
Minerva-7B-instruct-v1.0
Spanish
Salamandra-7B-instruct
TABLE IV: Models used for paraphrase generation.
PII type
Language
Model
BLEU
ROUGE-L
Sem. sim.
Email
English
Qwen2.5-7B-Instruct
23.51
49.21
83.43
French
Mistral-7B-Instruct-v0.3
8.21
15.02
69.23
Italian
Minerva-7B-instruct
22.08
39.94
46.99
Spanish
Salamandra-7B-instruct
6.62
18.84
41.35
German
DiscoLM_German_7b
41.89
56.41
68.05
German
Qwen2.5-7B-Instruct
47.80
67.70
90.89
TABLE V: Mean BLEU, ROUGE-L (F1), and bidirectional semantic similarity between original excerpts and their paraphrases, by PII type and language. Lower BLEU/ROUGE-L and higher semantic similarity indicate better paraphrase quality.
PII type
Model
BLEU
ROUGE-L
Sem. sim.
Email
DiscoLM_German_7b
41.89
56.41
68.05
Qwen2.5-7B-Instruct
47.80
67.70
90.89
Phone
DiscoLM_German_7b
49.31
62.21
69.85
Qwen2.5-7B-Instruct
41.98
63.94
90.70
Twitter
DiscoLM_German_7b
45.20
56.36
61.46
Qwen2.5-7B-Instruct
43.12
64.79
92.02
TABLE VI: Comparison of two candidate paraphrase models for German.
English
French
Italian
Spanish
German
Model
Strict
Paraphrase
Strict
Paraphrase
Strict
Paraphrase
Strict
Paraphrase
Strict
Paraphrase
Twitter
Llama 3.2 1B
340
68
255
56
243
16
260
36
213
25
Llama 3.2 3B
548
127
492
96
501
28
488
68
389
36
Qwen2.5 3B
468
146
395
104
372
19
384
55
298
31
Qwen2.5 7B
591
181
479
135
439
20
440
70
361
43
TABLE VII: Number of leaked PII instances for the strict translation and paraphrased attack prefixes. Results are reported for each model, language, and PII type.
Safety mechanisms for large language models (LLMs) remain predominantly English-centric, creating systematic vulnerabilities in multilingual deployment. Prior work shows that translating malicious prompts into other languages can substantially increase jailbreak success rates, exposing a structural cross-lingual security gap. We investigate whether such attacks can be mitigated through language-agnostic semantic similarity without retraining or language-specific adaptation. Our approach compares multilingual query embeddings against a fixed English codebook of jailbreak prompts, operating as a training-free external guardrail for black-box LLMs. We conduct a systematic evaluation across four languages, two translation pipelines, four safety benchmarks, three embedding models, and three target LLMs (Qwen, Llama, GPT-3.5). Our results reveal two distinct regimes of cross-lingual transfer. On curated benchmarks containing canonical jailbreak templates, semantic similarity generalizes reliably across languages, achieving near-perfect separability (AUC up to 0.99) and substantial reductions in absolute attack success rates under strict low-false-positive constraints. However, under distribution shift - on behaviorally diverse and heterogeneous unsafe benchmarks - separability degrades markedly (AUC ≈ 0.60-0.70), and recall in the security-critical low-FPR regime drops across all embedding models.
Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training. To study this phenomenon, we introduce STEER (Safety Targeted Embedding Exploit via Refinement), a gradient-guided attack that identifies words contributing most strongly to the model's refusal behavior and iteratively translates them into low-resource languages to suppress refusal while preserving harmful intent. Across six open-source 8B-parameter models, STEER achieves attack success rates of up to 93.0% on JailbreakBench and 96.7% on AdvBench, outperforming random code-switching and Greedy Coordinate Gradient (GCG). The resulting prompts also transfer to GPT-4o-mini, achieving a 35.5% attack success rate without requiring access to the target model, suggesting that the underlying weakness is not specific to a single architecture. These findings demonstrate that safety mechanisms aligned primarily on English cannot be assumed to generalize across multilingual inputs. We argue that improving multilingual safety requires broader coverage during alignment and mechanisms that explicitly detect and abstain on out-of-distribution inputs.
Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. However, unlearning research remains heavily English-centric. We study multilingual unlearning by extending the TOFU benchmark to five languages, and fine-tune, unlearn, and query our models with different permutations of languages. We find that unlearning transfer, the ability of an unlearned model to "forget" facts in languages other than the unlearning language, is highly variable: e.g., it is strongest between languages sharing scripts and families, and we show that the unlearning language predicts which query languages are most likely to yield the strongest transfer. Layer-wise analysis reveals that unlearning leaves the shared cross-lingual latent space largely intact in early layers, instead operating primarily in later decoding layers. This suggests that unlearning does not truly erase knowledge, but rather induces superficial suppression. Exploiting this structure, a single inference-time steering direction reverses much of this suppression across languages, recovering 50% (Qwen) and 90% (Gemma) of the unlearned knowledge.
Chaoyi Xiang, Olga Ohrimenko, Benjamin I. P. Rubinstein +1
School of Computing and Information Systems, The University of Melbourne, Melbourne, Australia.