Organizations: Harbin Institute of Technology, Shenzhen, China · Peng Cheng Lab, Shenzhen, China · Kuaishou, Beijing, China · Monash University, Melbourne, Australia
Instruction tuning aligns large language models (LLMs) with human intentions but requires diverse, high-quality data that are difficult to collect in privacy-sensitive domains. Federated instruction tuning (FedIT) enables collaborative training across data owners, yet existing methods typically assume sufficient local data. In realistic few-shot settings, limited samples can cause overfitting, degrade performance, and increase vulnerability to training data extraction attacks. We propose PPFedIT, a federated algorithm that improves both model performance and privacy protection in federated few-shot learning. It comprises three client-side steps: (1) synthetic data generation, which uses LLMs to diversify and enrich local data; (2) parameter isolation training, which updates the shared global LLM on synthetic data and local LLMs on private local data to mitigate synthetic-data noise; and (3) local aggregation then sharing, which mixes global and local model parameters before uploading them for server aggregation to mitigate data extraction attacks. Experiments on three open-source datasets show that PPFedIT improves model performance by an average of 8.4% and reduces the risk of data extraction attacks by approximately 20% in challenging federated few-shot settings.
Figures & tables
Figure 1. (a) Overview of the traditional FedIT framework and (b) the workflow of PPFedIT (illustrated for client k ). In contrast to FedIT , PPFedIT harnesses the generative power of the global LLM to expand and diversify the local data, boosting model utility in federated few-shot scenarios. Furthermore, to mitigate the notorious data extraction attacks that jeopardize the privacy of training data in FedIT , PPFedIT employs parameter-isolation training and local aggregation sharing, reinforcing client data safety throughout training.
Figure 2. The overview of PPFedIT . The innovation of PPFedIT compared to FedIT lies in client-side operations, including ① synthetic data generation, ② parameter isolation training, and ③ local aggregation sharing.
Data Statistics
Federated Setting
Datasets
|Train|
|Test|
Scenario
Rounds
|Clients|
Alpaca
1000
805
Cross-device
30
50
MedInstruct
500
216
Cross-silo
10
10
MedAlpaca
1000
400
Cross-device
30
50
Table 1. The data statistics and federated setting in our experiment.
Alpaca
MedInstruct
MedAlpaca
Avg.
Methods
Win ( ↑ )
Tie ( ↑ )
Lose ( ↓ )
Win ( ↑ )
Tie ( ↑ )
Lose ( ↓ )
Win ( ↑ )
Tie ( ↑ )
Lose ( ↓ )
Win ( ↑ )
Tie ( ↑ )
Lose ( ↓ )
Centralized
23.0
34.9
42.1
26.1
59.8
14.1
32.4
16.2
51.4
27.2
37.0
35.9
FedAvg
11.3
29.2
59.5
20.8
58.5
20.7
29.2
12.8
58.0
20.4
33.5
46.1
FedProx
12.3
30.4
57.3
22.5
58.0
19.5
28.5
13.2
58.3
21.1
33.9
45.0
SCAFFOLD
11.4
31.6
57.0
22.0
58.5
19.5
29.0
10.5
60.5
20.8
33.5
45.7
FedOPT
10.2
29.1
60.7
18.5
53.0
28.5
26.0
10.1
63.9
18.2
30.7
51.0
Table 2. The performance of PPFedIT and other contenders on federated few-shot instruction tuning. PPFedIT surpasses all federated algorithms and closely approaches the skyline algorithm Centralized performance.
Figure 3. The privacy data leakage measurement of Centralized , FedIT and PPFedIT during the training process. The light green and light red areas reflect the data leakage gaps between FedIT and Centralized , as well as FedIT and PPFedIT , in the same training stage.
Figure 4. (a) shows privacy data leakage of PPFedIT with different β and baselines proceeds with the federated training process. (b) shows the trade-off between privacy data leakage ( y -axis) and model utility ( x -axis). (c) shows the distribution of the Rouge-L similarity between private samples and their generated synthetic samples. We measure privacy data leakage using Rouge-L, where higher values indicate more significant privacy leakage. The model utility uses the WT score, representing the win and tie ratio sum. The parameter β in PPFedIT determines how much client privacy parameters are exposed. Pre-train refers to the backbone model without instruction tuning. PPFedIT demonstrates stronger privacy-preserving capabilities against training data extraction attacks than all baselines.
Figure 5. (a) The WT scores of various replaced synthetic data during the federated training process. (b) The contribution of FL to synthetic data generation.
Figure 6. The comparison of different synthetic data selection methods on MedInstruct (a) and MedAlpaca (b). Our IFS method more effectively screens higher-quality samples than the other methods.
Figure 7. (a) Convergence analysis of our method and the baseline methods on Alpaca and (b) Performances under different Non-IID Settings.
Figure 8. Impact of demonstration sample size (a) and synthetic data volume (b) during synthetic data generation.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9. The prompts used in our synthetic data generation. (a) Prompt used for generating new instructions. We randomly sample eight instructions from local data for in-context demonstration. The model is allowed to generate instructions for the new instruction. (b) Prompts used for input and output generation given the new instruction. We prompt the model with randomly sampled examples followed by the new instruction for the response generation.
Figure 10. GPT4-as-a-Judge prompt for evaluating the outputs of PPFedIT and baseline methods.
Quality Review Question
Alpaca
MedInstruct
MedAlpaca
Does the instruction describe a valid task?
91%
88%
85%
Is the response appropriate for the instruction and input?
52%
59%
57%
Is the correct format?
56%
66%
53%
Appendix
Table 3. Data quality review of the generated synthetic data.
Winogrande
ARC
Hellaswag
TruthfulQA
MMLU
Avg.
Pre-train
57.5
48.1
44.8
37.2
24.2
42.4
Ours
60.1
50.5
59.0
39.4
25.6
46.9
Appendix
Table 4. The performance of TinyLlama-1.1B on HuggingFace OpenLLM Benchmark using synthetic data of Alpaca . Pre-train refers to the backbone model without instruction tuning.
School of Artificial Intelligence, Beihang University Beijing, China · Center for the Applied Statistics, School of Statistics, Renmin University of China Beijing, China · School of Computer Science & Technology, Beijing Jiaotong University Beijing, China +2