Training generative machine learning models to produce synthetic tabular data has become a popular approach for enhancing privacy in data sharing. As this typically involves processing sensitive personal information, releasing either the trained model or generated synthetic datasets can still pose privacy risks. Yet, recent research, commercial deployments, and privacy regulations like the General Data Protection Regulation (GDPR) largely assess anonymity at the level of an individual dataset. In this paper, we rethink anonymity claims about synthetic data from a model-centric perspective, arguing that meaningful assessments must account for the underlying generative model and be grounded in state-of-the-art privacy attacks. This perspective better reflects real-world deployments, where trained models are often accessible for interaction or querying. We interpret the GDPR's definitions of personal data and anonymization under such access assumptions to identify the identifiability risks that must be mitigated and map them to privacy attacks across threat settings. We then argue that synthetic data techniques alone do not ensure sufficient anonymization. Finally, we compare the two mechanisms most commonly used with synthetic data -- Differential Privacy (DP) and Similarity-based Privacy Metrics (SBPMs) -- and argue that while DP can offer robust protections against identifiability risks, SBPMs lack adequate safeguards. Overall, our work connects regulatory notions of identifiability with model-centric privacy attacks, enabling more responsible and trustworthy assessment of synthetic data systems by researchers, practitioners, and policymakers.
Figures & tables
Figure 1 . Synthetic data generation using generative models.
Figure 2 . Privacy attacks on synthetic data: an adversary with black or white-box access to the model (and side information) extracts private information from the training data.
Figure 3 . Synthetic data generation with Differential Privacy (DP); model is trained while satisfying DP.
Figure 4 . Synthetic data generation with Similarity-based Privacy Metrics (SBPMs); the privacy of synthetic data is measured through SBPMs.
Regulatory Risk
Privacy Attacks Type
Specific Privacy Attack (Assumptions)
Singling out
Differencing Attacks
-
Linkability
Membership Inference Attacks
GroundHog ( Stadler et al., 2022 ) (black-box; training algo, representative data)
Querybased ( Houssiau et al., 2022a ) (black-box; training algo, representative data)
DOMIAS ( van Breugel et al., 2023 ) (no-box; reference test data)
AuditSynth ( Annamalai et al., 2024b ) (white-box; training algorithm)
MIDST ( Wu et al., 2025 ) (white/black-box; training algo, representative data)
Table 1 . Mapping between regulatory risks and privacy attack classes. The attacks operationalize particular manifestations of the corresponding regulatory risks, and the mappings are non-exhaustive.
Privacy Criterion
DP
SBPMs
Adversarial Model
well-defined
undefined
Privacy Guarantees
general
empirical
Privacy Analysis
worst-case
average-case
Privacy Risk
overestimation
underestimation
Plausible Deniability
yes
no
Privacy Subject
generative model
synthetic data
Table 2 . Comparison between synthetic data with privacy guaranteed by DP and SBPMs across eight privacy-related criteria.
Consideration
DP
SBPMs
Utility
reduced
not affected
Fairness
reduced/disparate
not affected
Consistency
not always
no
Interpretation
challenging
seems intuitive
Computational Performance
slow
fast
Implementation and Adoption
difficult
easy
Table 3 . Comparison between synthetic data with privacy guaranteed by DP and SBPMs across six non-privacy-related criteria.
Figure 5 . Base-case examples of synthetic data with privacy guaranteed by DP and SBPMs.
Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. In this position paper, we argue that this shift reflects an implicit change in community standards for what counts as sufficient privacy evidence, rather than a misunderstanding of well-established privacy principles. Drawing on an empirical analysis of recent publications across major ML venues, we show that synthetic data is frequently used in privacy-sensitive settings without explicit articulation of threat models, inference risks, or falsifiable privacy claims. As a result, privacy assurance often remains implicit, difficult to verify, and unevenly distributed, with heightened exposure for rare and minority records. We argue for treating privacy as an explicit, evidence-based scientific claim and recommend that ML venues adopt norms requiring privacy-relevant assertions to be clearly scoped, testable, and contestable.
Jiachen Zhao, Antonia Januszewicz, Taeho Jung
Department of Computer Science Engineering, University of Notre Dame, Notre Dame, United States.
The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. However, generating high-utility synthetic data often carries the risk of memorizing and regurgitating private information from the training corpus. In this work, we present a customizable empirical auditing framework designed to detect and explain such data disclosures. Our framework introduces a mechanism to distinguish between "true disclosures"-where the system directly reproduces a user's information-and "phantom disclosures''-where the system incidentally generates a user's data. By partitioning input data into training and holdout sets and applying rigorous statistical hypothesis testing, we determine if observed disclosures are consistent with strict privacy baselines, such as zero-learning or specific Differential Privacy (DP) bounds. Crucially, this approach requires no model access, no canary insertion, and no reference model training -only the synthetic output and a held-out control set. We demonstrate that this framework effectively functions as a membership inference attack, providing empirical lower bounds on privacy leakage that are tighter than prior data-based auditing methods. Our approach is model-agnostic, applies to any synthetic data generation mechanism, and requires orders of magnitude fewer computational resources than shadow-model or canary-based alternatives.
Tabular data sharing under privacy constraints is increasingly important for research and collaboration. Synthetic data generators (SDGs) are a promising solution, but synthetic data remains vulnerable to attacks, such as membership inference attacks (MIAs), which aim to determine whether a specific record was part of the training data. State-of-the-art MIAs are powerful but impractical: they rely on shadow modeling, requiring hundreds of SDG training runs, and need auxiliary data several times larger than the original training set. Fast proxy metrics like distance to closest record (DCR) are efficient but have limited sensitivity to MIA risk. We introduce ReMIA (Relative Membership Inference Attack), a practical privacy metric that requires only two SDG training runs and additional data no larger than the original training set. Rather than predicting whether a record was in the training set, ReMIA generates two synthetic datasets from two source datasets and measures whether a classifier can identify which source a record came from. Experiments across multiple tabular datasets and SDGs show that ReMIA has a sensitivity comparable to state-of-the-art MIAs while being substantially more practical. We further observe that SDGs can achieve privacy-utility trade-offs that traditional noise-based anonymization methods do not match. Code is available at https://github.com/aindo-com/remia.