As Large Language Models (LLMs) gain tool use capabilities such as web search, they can retrieve and cross-reference public information, creating privacy risks beyond memorization. One manifestation is re-identification: linking an anonymized interview transcript to a named individual. Yet without ground-truth identities, the coverage of such attacks and the protection offered by a defense cannot be reliably measured. We introduce AgentDOXX, an evaluation suite of 822 synthetic interview transcripts grounded in public information about real individuals with known identities. We evaluate fifteen configurations of open-weight and proprietary models, isolating the effect of web search, and analyze their search trajectories to distinguish retrieval-driven from parametric identifications. Ground-truth identities reveal that re-identification risk is distributed across an agent's execution: retrieval and parametric recall both contribute, with open-weight models identifying 15-28% of transcripts without search; identification succeeds in over 88% of cases once the target appears in a retrieved result; entity masking leaves at least one attacker successful on 85.3% of a stratified sample; and privacy instructions suppress naming but not retrieval, with configurations scoring 0% accuracy yet retrieving the subject in up to 62% of transcripts. We further show that observed attack trajectories can provide supervision for localizing identifying spans, offering a path toward attack-informed anonymization.
Figures & tables
Figure 1: Construction of AgentDoxx .
Predicted Real
Predicted Synthetic
Actual Real
103 (TP)
47 (FN)
Actual Synthetic
105 (FP)
45 (TN)
Table 1: Human judgments of real and synthetic interviews. Each of the 100 interviews receives three judgments.
Figure 2: Precision–recall and confidence calibration analysis.
Figure 3: Re-identification outcomes decomposed by the role of agentic web search versus parametric prior knowledge.
Model
Base
NER
Spans-redacted
NER+para
Spans-para
Noise
Privacy
DeepSeek V4 Flash
90.2
83.9
44.8
65.7
63.6
35.7
2.1
Gemini 3.1 Pro Preview
96.5
81.8
52.4
65.7
68.5
39.2
0.0
Gemini 3.7 Flash
97.2
81.1
55.9
62.9
68.5
44.1
50.3
GLM 5.3 Flash
97.2
79.0
53.8
64.3
66.4
51.7
24.5
GPT-4.1 Mini (+search)
42.0
26.6
11.2
18.9
16.8
15.4
38.5
GPT-4.1 Mini (no search)
11.2
5.6
3.5
4.9
4.2
5.6
6.3
Table 2: Re-identification accuracy (%) under six de-identification conditions ( n=143 transcripts).
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Answer level
Form of the abstention
Trajectory level
Model
Names
Abstains
Refusal
Empty
Subject retrieved
Abstains yet retrieved
of 143
of 143
of abstains
of abstains
of 143
of abstains
DeepSeek V4 Flash
3
140
54%
46%
26%
24%
Gemini 3.1 Pro Preview
4
139
99%
1%
62%
63%
Gemini 3.7 Flash
72
71
25%
75%
67%
37%
GLM 5.3 Flash
40
103
81%
19%
31%
8%
Appendix
Table 3: Model behaviour when instructed to protect privacy, across all 12×143=1716 model–transcript pairs. Answer level : whether the model’s final output names a specific individual (correctly or not), or abstains from naming anyone. Form of the abstention : how the model abstains. Refusal with analysis means the model explicitly states it is withholding the identity (e.g. Identity withheld (Privacy Constraint) ) but still describes the subject’s identifying attributes in a rationale field. Empty means the model returns identifier_candidates: [] with no explanation. These two columns sum to the abstention count. Trajectory level : whether the subject’s real name appeared somewhere in the search results the model retrieved during its reasoning process. Subject retrieved counts transcripts where the name appeared in at least one search result. Abstains yet retrieved counts cases where the model had already seen the correct name in its search results but still chose not to report it.
Figure 4: Mean attribute coverage by occupation for groups with n≥15 . Error bars denote standard errors of the mean.
Figure 5: Recall and precision of high-confidence candidates by occupation and model configuration. Each cell shows correct over total; shading encodes the ratio.
Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-identification, yet those same details also carry downstream analytic value of the text. Existing defenses either remove explicit identifiers, perturb text for formal privacy, or test rewritten text against non-web inference models, leaving underexplored the operating region between resistance to agentic web-search re-identification and utility retention. We introduce AURA (\textbf{A}nonymization with \textbf{U}tility-\textbf{R}etention \textbf{A}daptation), an LLM-powered \textit{mask-reconstruct} framework that decouples privacy localization from utility-preserving reconstruction and selects candidates with adversarial privacy and utility-retention checks. We evaluate AURA on real-user interview transcripts using re-identification attacks carried out by web-search agents, along with a utility evaluation based on interviewee-profile facts, codebook facts, and the joint contextual utility grid. Our results show that AURA improves the privacy-utility frontier by using adaptive privacy scope to strengthen resistance to agentic re-identification and using a mask-reconstruct anonymization method to better preserve contextual utility under fixed privacy scope.
Ziwen Li, Jianing Wen, Tianshi Li
Khoury College of Computer Sciences Northeastern University Boston, MA
As LLM-based agents increasingly browse the web on users' behalf, a natural question arises: can websites passively identify which underlying model powers an agent? Doing so would represent a significant security risk, enabling targeted attacks tailored to known model vulnerabilities. Across 14 frontier LLMs and four web environments spanning information retrieval and shopping tasks, we show that an agent's actions and interaction timings, captured via a passive JavaScript tracker, are sufficient to identify the underlying model with up to 96% F1. We formalise this attack surface by demonstrating that classifiers trained on agent actions generalise across model sizes and families. We further show that strong classifiers can be trained from few interaction traces and that agent identity can be inferred early within an episode. Injecting randomised timing delays between actions substantially degrades classifier performance, but does not provide robust protection: a classifier retrained on delayed traces largely recovers performance. We release our harness and a labelled corpus of agent traces here.
William Lugoloobi, Samuelle Marro, Jabez Magomere +2
Oxford Internet Institute, University of Oxford · Department of Engineering Science, University of Oxford
The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by the public. While prior research has established that mobility traces are highly unique and that individuals can, in principle, be identified from a handful of spatio-temporal points, such attacks have historically required significant manual effort from skilled analysts, limiting their practical scale. In this feasibility study, we demonstrate in a real world setting that agentic AI fundamentally changes this threat model. We present an end-to-end pipeline in which large language model agents autonomously search the open web, cross-reference public records and social media, and resolve raw coordinate sequences to candidate identities - without human intervention. We evaluate the pipeline on a spatio-temporal dataset containing simulated location points anchored at and around true home and work addresses, focusing on a high-risk disclosure scenario. Our results demonstrate that, from spatio-temporal data and public sources alone, our agentic AI successfully re-identified 18 of the 25 re-identifiable individuals (72%) and 18 of 43 cases overall (41.9%). We discuss implications for Statistical Disclosure Control (SDC) practice and outline the near-future escalation that data custodians and regulators must anticipate. De facto anonymity - an implicit foundation of SDC practice - is shifting. Agentic AI strengthens the case that re-identification is reasonably likely by any means under the GDPR Recital-26 standard, at costs of minutes-and-dollars per target.
Oscar Thees, Roman Müller, Matthias Templ
University of Applied Sciences and Arts Northwestern Switzerland (FHNW), Riggenbachstrasse 16, 4600 Olten, Switzerland