Organizations: Graduate School of Science and Technology, University of Tsukuba, 1-1-1, Tennodai, Tsukuba, 305-8573, Ibaraki, Japan · Institute of Systems and Information Engineering, University of Tsukuba, 1-1-1, Tennodai, Tsukuba, 305-8573, Ibaraki, Japan · Center for Artificial Intelligence Research, Tsukuba Institute for Advanced Research (TIAR), University of Tsukuba, 1-1-1, Tennodai, Tsukuba, 305-8577, Ibaraki, Japan
Geometric data perturbation enables one-shot representation sharing for privacy-preserving collaborative learning: each participant applies a secret distance-preserving transformation to its private data and uploads the resulting representation to a central analyst. We study analyst-participant collusion, in which a colluding participant discloses its data and transformation to help the analyst reconstruct another participant's data. Independent participant-specific transformations block direct inversion through a disclosed common transformation but leave uploads in incompatible coordinate systems, degrading pooled learning. Data Collaboration analysis restores compatibility by aligning transformed copies of a common anchor matrix withheld from the analyst. We show that, when the centered anchor matrix has full column rank, a colluder who discloses it enables exact recovery of every participant's transformation and inversion of noiseless private representations. Adding noise to private-data representations leaves this transformation-recovery channel intact and reduces leakage at a substantial utility cost. Instead, we perturb the anchor representations: each participant perturbs only its transformed anchor representation, preserving the geometry of its private-data upload while turning known-anchor transformation recovery into a noisy estimation problem. The analyst estimates the alignment using a spectral estimator for a generalized orthogonal Procrustes problem. We analyze recovery attacks against this protocol and compare both noise placements on the CelebA and VGGFace2 facial image datasets. Under the evaluated collusion attacks, noisy-anchor alignment retains higher downstream accuracy at low identity-linkage levels. Participant-count experiments examine the utility gains and limitations of larger collaborations at comparable measured linkage.
Figures & tables
Fig. 1: Overview of the NAA-GDP pipeline.
Method
Private-data upload Yi
Anchor upload Bi
Local
—
—
C-GDP
XiOs+1niΨs⊤
—
C-GDP (private-data noise)
XiOs+1niΨs⊤+σWi
—
I-GDP
XiOi+1niΨi⊤
—
NAA-GDP
XiOi+1niΨi⊤
AOi+1rΨi⊤+vWi
TABLE I: Compared methods and their uploads
Stage
Output dimensions
Reshape
1×20×20
Convolution 1, ReLU
10×16×16
Max pooling 1
10×8×8
Convolution 2, ReLU
20×4×4
Max pooling 2
20×2×2
Flatten
80
TABLE II: CNN architecture for the CelebA and VGGFace2 datasets.
Fig. 2: Privacy–utility trade-off on the CelebA (left) and VGGFace2 (right) datasets. Bars show mean test balanced accuracy; error bars show one sample standard deviation.
Fig. 3: Visual reconstructions under analyst-participant collusion on the CelebA (a) and VGGFace2 (b) datasets. Each column pairs a query with a different reference photograph of the same identity. C-GDP provides the noiseless decoded baseline. Percentages are condition-wide linkage accuracies over 1,000 lineups, not individual-image scores; shared column positions do not imply matched linkage between noise families.
Fig. 4: Participant-count sensitivity of anchor-noise NAA-GDP on the CelebA (left) and VGGFace2 (right) datasets. Points show mean utility within each linkage interval; error bars show one sample standard deviation.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Fig. 5: Identity linkage under MP, AM, and OP reconstruction attacks on NAA-GDP: the CelebA (left) and VGGFace2 (right) datasets. The dashed line marks 10% random-guessing accuracy.
Collaborative analysis of decentralized confidential datasets is important, but direct sharing of original datasets is often restricted by privacy and institutional constraints. Data collaboration (DC) analysis transforms each dataset into privacy-preserving intermediate representations via party-specific obfuscation functions and integrates them into common collaboration representations using an anchor dataset. However, many existing DC analysis methods rely on linear transformations for data obfuscation and integration, which may increase reconstruction risk. Although nonlinear dimensionality reduction can mitigate this risk, conventional linear integration methods cannot accurately align intermediate representations produced by nonlinear transformations. Moreover, existing integration methods mainly minimize discrepancies among parties and do not explicitly incorporate geometric or target-variable information useful for downstream analysis. To overcome these limitations, we first formulate linear kernel integration (LKI) as a linear integration method and then kernelize it to obtain nonlinear kernel integration (NKI). NKI admits a globally optimal solution via kernel ridge regression and an eigenvalue problem. We also introduce graph regularization and a centering constraint so that the target representation can capture geometric and target-variable information useful for downstream analysis. Experiments on image classification tasks demonstrate that NKI improves classification accuracy over existing linear integration methods under nonlinear dimensionality reduction, with further gains from target-variable-aware graph regularization and centering. The results also show that dimensionality reduction choices substantially affect both classification accuracy and reconstruction risk.
Yamato Suetake, Yuta Kawakami, Shunnosuke Ikeda +1
Graduate School of Science and Technology, University of Tsukuba, Tsukuba, Ibaraki, Japan · Institute of Systems and Information Engineering, University of Tsukuba, Tsukuba, Ibaraki, Japan
Federated Learning (FL) enables collaborative model training among multiple parties without centralizing raw data. There are two main paradigms in FL: Horizontal FL (HFL), where all participants share the same feature space but hold different samples, and Vertical FL (VFL), where parties possess complementary features for the same set of samples. A prerequisite for VFL training is privacy-preserving entity alignment (PPEA), which establishes a common index of samples across parties (alignment) without revealing which samples are shared between them. Conventional private set intersection (PSI) achieves alignment but leaks intersection membership, exposing sensitive relationships between datasets. The standard private set union (PSU) mitigates this risk by aligning on the union of identifiers rather than the intersection. However, existing approaches are often limited to two parties or lack support for typo-tolerant matching. In this paper, we introduce the Sherpa.ai multi-party PSU protocol for VFL, a PPEA method that hides intersection membership and enables both exact and noisy matching. The protocol generalizes two-party approaches to multiple parties with low communication overhead and offers two variants: an order-preserving version for exact alignment and an unordered version tolerant to typographical and formatting discrepancies. We prove correctness and privacy, analyze communication and computational (exponentiation) complexity, and formalize a universal index mapping from local records to a shared index space. This multi-party PSU offers a scalable, mathematically grounded protocol for PPEA in real-world VFL deployments, such as multi-institutional healthcare disease detection, collaborative risk modeling between banks and insurers, and cross-domain fraud detection between telecommunications and financial institutions, while preserving intersection privacy.
Daniel M. Jimenez-Gutierrez, Dario Pighin, Enrique Zuazua +4
We propose a decentralized privacy-preserving learning algorithm in which each agent holds a single private sample and a shared model. Samples are learned sequentially, and each update must preserve the endpoint mappings at previously learned samples while protecting private data. This gives each agent three roles: (i) a learner that updates the model parameters, (ii) a teacher whose sample is learned at the current iteration, and (iii) a protected agent whose sample has already been learned. We build on Tuning without Forgetting (TwF) method to preserve previously learned mappings and show that TwF provides an indistinguishability guarantee for the learner whenever the set of protected agents contains another sample with the same label. For the teacher, we formulate a minimax optimal control problem that models the differential privacy noise as a worst-case disturbance to prevent performance loss while maintaining the same level of privacy for the gradient. For the protected agents, we compute the projections locally and aggregate them using a private push-sum gossip protocol. We prove geometric convergence of the decentralized gossip algorithm and of the distributed projection for TwF.