Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning
Authors: Quan Minh Nguyen, Min-Seon Kim, Hoang M. Ngo, Trong Nghia Hoang, Hyuk-Yoon Kwon, My T. Thai
Organizations: CISE, University of Florida, FL, USA · North Carolina State University, NC, USA · Washington State University, WA, USA · Seoul National University of Science and Technology, South Korea
Membership inference attacks (MIAs) pose a serious privacy threat in federated learning (FL). While MIAs have been extensively studied in standard FL, the recent shift toward federated fine-tuning introduces new and largely unexplored attack surfaces. In this work, we show that federated prompt-tuning, which adapts pre-trained foundation models using lightweight input prefixes, exposes a novel and effective vector for membership inference. We propose PromptMIA, a membership inference attack tailored to federated prompt-tuning, in which a malicious server introduces adversarially crafted prompts and exploits their updates during collaborative training to determine whether a target data point belongs to a client's private dataset. We formalize this threat via a security game and demonstrate that PromptMIA achieves consistently high attack advantage across diverse benchmark datasets, substantially outperforming current SOTA federated MIAs. We also provide a theoretical lower bound on the attack advantage that explains the observed empirical behavior. Finally, we show that existing MIA defenses are often ineffective against PromptMIA, highlighting the need for defense mechanisms specifically tailored to prompt-tuning in federated settings.
Figures & tables
Figure 1 : PromptMIA workflow: (1) the server injects adversarial prompts designed for a target sample into the global prompt pool; (2) the modified pool is broadcast to clients; (3) each client performs query-key matching, which selects all adversarial prompts if the target sample is present; (4) selected prompts are locally updated and (5) returned to the server. By monitoring which prompts are updated, the server infers the target’s membership in client data.
Figure 2 : Comparison of the global key distributions produced by a ViT-B/32 model trained on CIFAR-10 after 60 global epochs, visualized using t-SNE. Blue keys are benign keys, and red keys are adversarial keys.
Figure 3 : Left: t-SNE projection of the global keys and query vectors from the train set. Center: K-Means clustering of keys with queries. Right: Gaussian Mixture Query Model. Visualizations are generated from a ViT-B32 model trained on CIFAR-10 for 60 global epochs.
Figure 4 : Performance of PromptMIA vs Naive averaged across three models . Each subplot shows Advantage and A.S.R w.r.t Batch Size across CIFAR10, CIFAR100, TinyImageNet, and FourDataset.
Figure 5Table 6
Figure 7 : Attack Success Rate of PromptMIA under Input Noise Perturbation with different ϵ .
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : (Local Phase) each client samples and fine-tunes a subset of global summarizing prompts using a prompt-selection strategy; (Global Phase) the server aggregates all local prompt sets to refine the global prompt pool.
#
Attack
Attack Surface
Attack Methodology
Assumptions
1
Blackbox-Loss [ 56 ] Loss-Series [ 17 ]
Data Loss
Model tends to have lower loss on training points than on non-training points. [2] uses loss values over multiple rounds.
Bounded loss and Gaussian error distributions
2
Grad-Cosine [ 28 ]
Gradient vector
The distributions of the cosine similarity of instances from the training data (members) and from non-members are different.
Gradient vectors of different instances are orthogonal & their cosine similarity is Gaussian.
3
FedMIA [ 58 ]
Data Loss (Fed-MIA-I), Gradient vector (Fed-MIA-II)
Update distribution of clients trained with target data is different from non-target clients
Update distribution follows Gaussian distribution
4
PromptMIA (ours)
Global prompt pool in FPT
Attacker injects adversarial prompts that will always be updated when target data is present.
Each query vector is modeled as a Gaussian mixture over benign keys
Appendix
Table 3: Comparison of PromptMIA and existing Membership Inference Attacks.
Figure 9 : Impact of Prompt Dropout on attack success rate.
Figure 10 : Gaussian Mixture Query Model under extreme Non-IID.
Figure 11 : Comparison of theoretical FPR bounds under varying non-IID conditions.
Figure 12 : PromptMIA vs Naive attack results against ViT-B/32. Each subplot shows Advantage and Attack Success rate w.r.t Batch Size across CIFAR10, CIFAR100, TinyImageNet, and FourDataset.
Figure 13 : PromptMIA vs Naive attack results against ConViT. Each subplot shows Advantage and Attack Success rate w.r.t Batch Size across CIFAR10, CIFAR100, TinyImageNet, and FourDataset.
Figure 14 : PromptMIA vs Naive attack results against Deit. Each subplot shows Advantage and Attack Success rate w.r.t Batch Size across CIFAR10, CIFAR100, TinyImageNet, and FourDataset.
Figure 15 : Visualization of outlier detection methods on CIFAR-10 trained Vit-B32. Blue keys are benign keys. Red keys are adversarial keys. Crossed keys are flagged as outliers from the corresponding algorithm.
Figure 16 : Visualization of outlier detection methods on CIFAR-10 trained ViT-B32. Blue keys are benign keys. Red keys are adversarial keys. Crossed keys are flagged as outliers from the corresponding algorithm. Outlier detection methods still falsely flag benign keys as outliers when no adversarial keys are present.
Dataset
Model
Method
Precision
Recall
F1
CIFAR10
ViT
IsolationForest
0.2828
1.0000
0.4409
LocalOutlierFactor
0.0000
0.0000
0.0000
OneClassSVM
0.3052
0.4773
0.3723
EllipticEnvelope
0.0000
0.0000
0.0000
ConViT
IsolationForest
0.2330
1.0000
0.3779
LocalOutlierFactor
0.0000
0.0000
0.0000
Appendix
Table 4 : Precision, Recall, and F1 of Outlier Detection across datasets, models, and methods.
Figure 17 : Ablation study on PromptMIA . Each subfigure shows the effect of one parameter: (a) M , (b) N , (c) δmin , (d) Δ and (e) training rounds.
Figure 18 : Visualization of distribution of benign keys (blue) and query vectors q(x) (green) across training rounds.
Figure 19 : Attack success rate of PromptMIA under increasing batch size ( up to 1024) across different datasets. The results consistently show that the attack success rate remains high even under extremely large batch size.
Figure 20 : Attack Success Rate of PROMPTMIA under multimodal and text input modality.
Figure 21 : Attack success rate of PROMPTMIA against prompt-based FL on CIFAR-10 under different heterogeneity settings. Even on extremely large batch size, the attack success rate remains highly significant at more than 85%.
Figure 22 : Visualization of prompt clusters produced by PFPT when trained on CIFAR10 with various heterogeneous settings
Figure 23 : Impact of number of clients on PromptMIA .
Dataset
Model Accuracy
Model Accuracy With PromptMIA
CIFAR10
0.95(0.01)
0.94(0.01)
CIFAR100
0.78(0.01)
0.76(0.02)
Appendix
Table 5 : Impact of prompt injection on model accuracy across different datasets.
Model
Total params
Modified params
Sparsity (%)
vit_b32
89,362,418
3,072
0.0034
deit_b16
86,818,058
3,072
0.0035
convit_base
86,790,442
3,072
0.0035
Appendix
Table 6 : Attack Sparsity of PromptMIA .
Metric
Adversarial key generation
FLOPs
86,000 (86 KFLOPs)
Latency
351 μ s ±7.7 (0.351 ms)
Throughput
11385±252 keys/sec
Appendix
Table 7 : Computational overhead for adversarial key generation.
Federated Large Language Models (FedLLMs) enable multiple parties to collaboratively fine-tune LLMs without sharing raw data, addressing challenges of limited resources and privacy concerns. Despite data localization, shared gradients can still expose sensitive information through membership inference attacks (MIAs). However, FedLLMs' unique properties, i.e. massive parameter scales, rapid convergence, and sparse, non-orthogonal gradients, render existing MIAs ineffective. To address this gap, we propose ProjRes, the first projection residuals-based passive MIA tailored for FedLLMs. ProjRes leverages hidden embedding vectors as sample representations and analyzes their projection residuals on the gradient subspace to uncover the intrinsic link between gradients and inputs. It requires no shadow models, auxiliary classifiers, or historical updates, ensuring efficiency and robustness. Experiments on four benchmarks and four LLMs show that ProjRes achieves near 100% accuracy, outperforming prior methods by up to 75.75%, and remains effective even under strong differential privacy defenses. Our findings reveal a previously overlooked privacy vulnerability in FedLLMs and call for a re-examination of their security assumptions. Our code and data are available at [link](https://anonymous.4open.science/r/Passive−MIA−5268).
Guilin Deng, Silong Chen, Yuchuan Luo +6
Colleague of Computer Science and Technology, National University of Defense Technology, Changsha, China · City University of Hong Kong, Hong Kong, China · Shenzhen University, Shenzhen, China
Instruction tuning aligns large language models (LLMs) with human intentions but requires diverse, high-quality data that are difficult to collect in privacy-sensitive domains. Federated instruction tuning (FedIT) enables collaborative training across data owners, yet existing methods typically assume sufficient local data. In realistic few-shot settings, limited samples can cause overfitting, degrade performance, and increase vulnerability to training data extraction attacks. We propose PPFedIT, a federated algorithm that improves both model performance and privacy protection in federated few-shot learning. It comprises three client-side steps: (1) synthetic data generation, which uses LLMs to diversify and enrich local data; (2) parameter isolation training, which updates the shared global LLM on synthetic data and local LLMs on private local data to mitigate synthetic-data noise; and (3) local aggregation then sharing, which mixes global and local model parameters before uploading them for server aggregation to mitigate data extraction attacks. Experiments on three open-source datasets show that PPFedIT improves model performance by an average of 8.4% and reduces the risk of data extraction attacks by approximately 20% in challenging federated few-shot settings.
Zhuo Zhang, Jingyuan Zhang, Jintao Huang +6
Harbin Institute of Technology, Shenzhen, China · Peng Cheng Lab, Shenzhen, China · Kuaishou, Beijing, China +1
Federated prompt tuning (FPT) enables collaborative adaptation of vision--language models (VLMs) using lightweight prompts. Existing methods often address heterogeneity and privacy through a split-prompt design under local differential privacy (DP), combining a shared prompt for global transfer with private prompts for local adaptation. However, a single shared prompt may over-smooth diverse transferable knowledge, weakening the balance between personalization and generalization. Multi-expert prompts (MEPs) can better capture this diversity, but enlarge the communicated space, increasing DP noise and communication cost while making robust expert composition more difficult. We propose FedSEPT, a privacy-preserving Fed}erated Subspace-decomposed Expert Prompt Tuning. Specifically, we employ Subspace-decomposed Expert Modeling (SEM) to parameterize multiple prompt experts with shared low-rank factors, a fixed public basis, and private residuals, thereby confining communication and DP perturbation to a compact factor space while enabling direct server aggregation in a common coordinate system. We further design Instance-aware Expert Fusion (IEF), which adaptively combines semantically complementary experts via on-device routing and performs efficient logit-level fusion using cached expert-specific text features. Extensive experiments on 11 heterogeneous benchmarks show that, under the same privacy constraints, FedSEPT achieves a better trade-off between local adaptation and global generalization than strong baselines.
Yuhua Wang, Xiaodong Li, Yihao Guo +6
School of Artificial Intelligence, Beihang University Beijing, China · Center for the Applied Statistics, School of Statistics, Renmin University of China Beijing, China · School of Computer Science & Technology, Beijing Jiaotong University Beijing, China +2