cs.CROct 4, 2026

AgentDoxx: Agentic Re-identification of Anonymized Text with Web Search

Authors: Jianing Wen, Tianshi Li

Organizations: Khoury College of Computer Sciences Northeastern University Boston, MA

Abstract

As Large Language Models (LLMs) gain tool use capabilities such as web search, they can retrieve and cross-reference public information, creating privacy risks beyond memorization. One manifestation is re-identification: linking an anonymized interview transcript to a named individual. Yet without ground-truth identities, the coverage of such attacks and the protection offered by a defense cannot be reliably measured. We introduce AgentDOXX, an evaluation suite of 822 synthetic interview transcripts grounded in public information about real individuals with known identities. We evaluate fifteen configurations of open-weight and proprietary models, isolating the effect of web search, and analyze their search trajectories to distinguish retrieval-driven from parametric identifications. Ground-truth identities reveal that re-identification risk is distributed across an agent's execution: retrieval and parametric recall both contribute, with open-weight models identifying 15-28% of transcripts without search; identification succeeds in over 88% of cases once the target appears in a retrieved result; entity masking leaves at least one attacker successful on 85.3% of a stratified sample; and privacy instructions suppress naming but not retrieval, with configurations scoring 0% accuracy yet retrieving the subject in up to 62% of transcripts. We further show that observed attack trajectories can provide supervision for localizing identifying spans, offering a path toward attack-informed anonymization.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LLM Anonymization Against Agentic Re-Identification

    May 29, 2026Ziwen Li, Jianing Wen, Tianshi LiAnonymizationAttacker Large Language Model

  2. Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces

    May 14, 2026William Lugoloobi, Samuelle Marro, Jabez Magomere +2Attacker Large Language ModelFingerprint

  3. Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy

    Jun 26, 2026Oscar Thees, Roman Müller, Matthias TemplRe-IdentificationMobility