cs.CRSep 27, 2026

The Privacy Fallacy of Crowdsourced Fine-Tuning: Extracting Proprietary Data via Topic-Based Poisoning

Authors: Sae Furukawa, Alina Oprea

Organizations: Khoury College of Computer Sciences, Northeastern University

Abstract

Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks. Crowdsourcing user conversations is an established approach to collecting SFT data at scale while reducing the need for costly manual annotation. However, it also allows untrusted users to contribute data to the fine-tuning pipeline. We investigate an underexplored privacy risk arising from this setting: can a malicious user poison a small fraction of the crowdsourced data to amplify extraction of previously unseen instructions contributed by other users? We show that this is possible using only black-box, output-only access to the deployed model. Experiments across four models and two datasets demonstrate substantial increases in training-data extraction: with only 50 poisoned examples, near-verbatim extraction reaches 3.71×3.71\times the rate without poisoning for Qwen2.5-14B on OpenMathInstruct and 3.08×3.08\times for Llama-3.1-8B on AceReason. Data filtering also proves largely ineffective in detecting poisoned samples: even the best-performing method achieves only 0.378 in F-1 score, leaving the majority of poisoned samples undetected. These findings demonstrate that seemingly benign crowdsourced contributions can amplify leakage of other records while remaining difficult to identify through data filtering.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Loss Landscape Poisoning: Targeted Extraction of Unseen Training Data from LLMs

    Jun 15, 2026Md Abdullah Al Mamun, Ngoc Phu Doan, Pedram Zaree +2Attacker Large Language ModelPoisoning

  2. Advancing the State-of-the-Art in Empirical Privacy Auditing

    Jun 9, 2026Nicole Mitchell, Galen Andrew, Arun Ganesh +2PrivacyModel Auditing

  3. DP-SelFT: Differentially Private Selective Fine-Tuning for Large Language Models

    May 17, 2026Haichao Sha, Zihao Wang, Yuncheng Wu +2Large Language Model Fine-TuningContinual Fine-Tuning