Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning
Authors: Quan Minh Nguyen, Min-Seon Kim, Hoang M. Ngo, Trong Nghia Hoang, Hyuk-Yoon Kwon, My T. Thai
Organizations: CISE, University of Florida, FL, USA · North Carolina State University, NC, USA · Washington State University, WA, USA · Seoul National University of Science and Technology, South Korea
Membership inference attacks (MIAs) pose a serious privacy threat in federated learning (FL). While MIAs have been extensively studied in standard FL, the recent shift toward federated fine-tuning introduces new and largely unexplored attack surfaces. In this work, we show that federated prompt-tuning, which adapts pre-trained foundation models using lightweight input prefixes, exposes a novel and effective vector for membership inference. We propose PromptMIA, a membership inference attack tailored to federated prompt-tuning, in which a malicious server introduces adversarially crafted prompts and exploits their updates during collaborative training to determine whether a target data point belongs to a client's private dataset. We formalize this threat via a security game and demonstrate that PromptMIA achieves consistently high attack advantage across diverse benchmark datasets, substantially outperforming current SOTA federated MIAs. We also provide a theoretical lower bound on the attack advantage that explains the observed empirical behavior. Finally, we show that existing MIA defenses are often ineffective against PromptMIA, highlighting the need for defense mechanisms specifically tailored to prompt-tuning in federated settings.
Figures & tables
Figure 1 : PromptMIA workflow: (1) the server injects adversarial prompts designed for a target sample into the global prompt pool; (2) the modified pool is broadcast to clients; (3) each client performs query-key matching, which selects all adversarial prompts if the target sample is present; (4) selected prompts are locally updated and (5) returned to the server. By monitoring which prompts are updated, the server infers the target’s membership in client data.
Figure 2 : Comparison of the global key distributions produced by a ViT-B/32 model trained on CIFAR-10 after 60 global epochs, visualized using t-SNE. Blue keys are benign keys, and red keys are adversarial keys.
Figure 3 : Left: t-SNE projection of the global keys and query vectors from the train set. Center: K-Means clustering of keys with queries. Right: Gaussian Mixture Query Model. Visualizations are generated from a ViT-B32 model trained on CIFAR-10 for 60 global epochs.
Figure 4 : Performance of PromptMIA vs Naive averaged across three models . Each subplot shows Advantage and A.S.R w.r.t Batch Size across CIFAR10, CIFAR100, TinyImageNet, and FourDataset.
Figure 5Table 6
Figure 7 : Attack Success Rate of PromptMIA under Input Noise Perturbation with different ϵ .
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : (Local Phase) each client samples and fine-tunes a subset of global summarizing prompts using a prompt-selection strategy; (Global Phase) the server aggregates all local prompt sets to refine the global prompt pool.
#
Attack
Attack Surface
Attack Methodology
Assumptions
1
Blackbox-Loss [ 56 ] Loss-Series [ 17 ]
Data Loss
Model tends to have lower loss on training points than on non-training points. [2] uses loss values over multiple rounds.
Bounded loss and Gaussian error distributions
2
Grad-Cosine [ 28 ]
Gradient vector
The distributions of the cosine similarity of instances from the training data (members) and from non-members are different.
Gradient vectors of different instances are orthogonal & their cosine similarity is Gaussian.
3
FedMIA [ 58 ]
Data Loss (Fed-MIA-I), Gradient vector (Fed-MIA-II)
Update distribution of clients trained with target data is different from non-target clients
Update distribution follows Gaussian distribution
4
PromptMIA (ours)
Global prompt pool in FPT
Attacker injects adversarial prompts that will always be updated when target data is present.
Each query vector is modeled as a Gaussian mixture over benign keys
Appendix
Table 3: Comparison of PromptMIA and existing Membership Inference Attacks.
Figure 9 : Impact of Prompt Dropout on attack success rate.
Figure 10 : Gaussian Mixture Query Model under extreme Non-IID.
Figure 11 : Comparison of theoretical FPR bounds under varying non-IID conditions.
Figure 12 : PromptMIA vs Naive attack results against ViT-B/32. Each subplot shows Advantage and Attack Success rate w.r.t Batch Size across CIFAR10, CIFAR100, TinyImageNet, and FourDataset.
Figure 13 : PromptMIA vs Naive attack results against ConViT. Each subplot shows Advantage and Attack Success rate w.r.t Batch Size across CIFAR10, CIFAR100, TinyImageNet, and FourDataset.
Figure 14 : PromptMIA vs Naive attack results against Deit. Each subplot shows Advantage and Attack Success rate w.r.t Batch Size across CIFAR10, CIFAR100, TinyImageNet, and FourDataset.
Figure 15 : Visualization of outlier detection methods on CIFAR-10 trained Vit-B32. Blue keys are benign keys. Red keys are adversarial keys. Crossed keys are flagged as outliers from the corresponding algorithm.
Figure 16 : Visualization of outlier detection methods on CIFAR-10 trained ViT-B32. Blue keys are benign keys. Red keys are adversarial keys. Crossed keys are flagged as outliers from the corresponding algorithm. Outlier detection methods still falsely flag benign keys as outliers when no adversarial keys are present.
Dataset
Model
Method
Precision
Recall
F1
CIFAR10
ViT
IsolationForest
0.2828
1.0000
0.4409
LocalOutlierFactor
0.0000
0.0000
0.0000
OneClassSVM
0.3052
0.4773
0.3723
EllipticEnvelope
0.0000
0.0000
0.0000
ConViT
IsolationForest
0.2330
1.0000
0.3779
LocalOutlierFactor
0.0000
0.0000
0.0000
Appendix
Table 4 : Precision, Recall, and F1 of Outlier Detection across datasets, models, and methods.
Figure 17 : Ablation study on PromptMIA . Each subfigure shows the effect of one parameter: (a) M , (b) N , (c) δmin , (d) Δ and (e) training rounds.
Figure 18 : Visualization of distribution of benign keys (blue) and query vectors q(x) (green) across training rounds.
Figure 19 : Attack success rate of PromptMIA under increasing batch size ( up to 1024) across different datasets. The results consistently show that the attack success rate remains high even under extremely large batch size.
Figure 20 : Attack Success Rate of PROMPTMIA under multimodal and text input modality.
Figure 21 : Attack success rate of PROMPTMIA against prompt-based FL on CIFAR-10 under different heterogeneity settings. Even on extremely large batch size, the attack success rate remains highly significant at more than 85%.
Figure 22 : Visualization of prompt clusters produced by PFPT when trained on CIFAR10 with various heterogeneous settings
Figure 23 : Impact of number of clients on PromptMIA .
Dataset
Model Accuracy
Model Accuracy With PromptMIA
CIFAR10
0.95(0.01)
0.94(0.01)
CIFAR100
0.78(0.01)
0.76(0.02)
Appendix
Table 5 : Impact of prompt injection on model accuracy across different datasets.
Model
Total params
Modified params
Sparsity (%)
vit_b32
89,362,418
3,072
0.0034
deit_b16
86,818,058
3,072
0.0035
convit_base
86,790,442
3,072
0.0035
Appendix
Table 6 : Attack Sparsity of PromptMIA .
Metric
Adversarial key generation
FLOPs
86,000 (86 KFLOPs)
Latency
351 μ s ±7.7 (0.351 ms)
Throughput
11385±252 keys/sec
Appendix
Table 7 : Computational overhead for adversarial key generation.
Colleague of Computer Science and Technology, National University of Defense Technology, Changsha, China · City University of Hong Kong, Hong Kong, China · Shenzhen University, Shenzhen, China
School of Artificial Intelligence, Beihang University Beijing, China · Center for the Applied Statistics, School of Statistics, Renmin University of China Beijing, China · School of Computer Science & Technology, Beijing Jiaotong University Beijing, China +2