cs.CRMar 10, 2024

PPFedIT: Towards Privacy-Preserving Federated Instruction Tuning with Few-shot Local Examples

Authors: Zhuo Zhang, Jingyuan Zhang, Jintao Huang, Hui Wang, Yue Yu, Hongzhi Zhang, Xun Zhou, Lizhen Qu, +1 more

Organizations: Harbin Institute of Technology, Shenzhen, China · Peng Cheng Lab, Shenzhen, China · Kuaishou, Beijing, China · Monash University, Melbourne, Australia

Abstract

Instruction tuning aligns large language models (LLMs) with human intentions but requires diverse, high-quality data that are difficult to collect in privacy-sensitive domains. Federated instruction tuning (FedIT) enables collaborative training across data owners, yet existing methods typically assume sufficient local data. In realistic few-shot settings, limited samples can cause overfitting, degrade performance, and increase vulnerability to training data extraction attacks. We propose PPFedIT, a federated algorithm that improves both model performance and privacy protection in federated few-shot learning. It comprises three client-side steps: (1) synthetic data generation, which uses LLMs to diversify and enrich local data; (2) parameter isolation training, which updates the shared global LLM on synthetic data and local LLMs on private local data to mitigate synthetic-data noise; and (3) local aggregation then sharing, which mixes global and local model parameters before uploading them for server aggregation to mitigate data extraction attacks. Experiments on three open-source datasets show that PPFedIT improves model performance by an average of 8.4% and reduces the risk of data extraction attacks by approximately 20% in challenging federated few-shot settings.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Privacy-Preserving Split Learning for Federated LLM Fine-Tuning

    Sep 9, 2026Heng Jin, Chaoyu Zhang, Hexuan Yu +2Federated LearningLarge Language Model Fine-Tuning

  2. Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning

    Jan 10, 2026Quan Minh Nguyen, Min-Seon Kim, Hoang M. Ngo +3Federated LearningMembership Inference Attacks

  3. Towards Privacy-Preserving Federated Prompt Tuning under Data Heterogeneity: A Subspace-Decomposed Expert Approach

    Jul 23, 2026Yuhua Wang, Xiaodong Li, Yihao Guo +6Soft Prompt TuningHeterogeneity