cs.CLOct 5, 2026

Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PII

Authors: Alexandru Nazare, Agnese Profico, Nicolò Vania, Elena Di Croce, Daria Caramanica, Davide Venditti, Elena Sofia Ruzzetti, Giancarlo A. Xompero, +1 more

Organizations: Human-Centric ART, University of Rome Tor Vergata · Department of Computer Science, University of Luxembourg · Almawave Labs, Rome

Abstract

The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (TDE) attacks when prompted in non-English languages. We construct a multi-domain PII dataset comprising social media handles, email addresses, and phone numbers and translate the attack contexts into Italian, Spanish, French, and German. Our results show that TDE attacks against both English-centric and multilingual models transfer to different languages: the attacks are successful on translated prompts, even though only the original English prompt might have been included in the pre-training data. A web-presence check on a sample of the translations confirms that they are not available online. The share of English leaks recovered in other languages grows with the multilingual capability of the model, and it drops sharply when the original wording is lost, even without a change of language. This suggests that native multilingual pre-training facilitates the emergence of latent cross-linguistic bridges that simplify the retrieval of personally identifiable information (PII). We analyze the activations of multilingual large language models (LLMs) and find that different translations of the same prompt are bridged in similar representations, with the strongest alignment in the middle layers. Our results highlight a fundamental security gap in modern LLMs, necessitating more robust, language-agnostic sanitization strategies for future model alignment.

Figures & tables

Explore similar work

CardsList
  1. Cross-Lingual Jailbreak Detection via Semantic Codebooks

    Apr 28, 2026Shirin Alanova, Bogdan Minko, Sabrina Sadiekh +1Multilingual Language Model EvaluationLLM Security

  2. Safety Targeted Embedding Exploit via Refinement

    Jul 2, 2026Joshua Adrian CahyonoMultilingual Language Model EvaluationAdversarial Attacks on LLMs

  3. Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility

    Jun 2, 2026Chaoyi Xiang, Olga Ohrimenko, Benjamin I. P. Rubinstein +1Machine UnlearningLLM Unlearning