The Privacy Fallacy of Crowdsourced Fine-Tuning: Extracting Proprietary Data via Topic-Based Poisoning
Organizations: Khoury College of Computer Sciences, Northeastern University
Abstract
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks. Crowdsourcing user conversations is an established approach to collecting SFT data at scale while reducing the need for costly manual annotation. However, it also allows untrusted users to contribute data to the fine-tuning pipeline. We investigate an underexplored privacy risk arising from this setting: can a malicious user poison a small fraction of the crowdsourced data to amplify extraction of previously unseen instructions contributed by other users? We show that this is possible using only black-box, output-only access to the deployed model. Experiments across four models and two datasets demonstrate substantial increases in training-data extraction: with only 50 poisoned examples, near-verbatim extraction reaches the rate without poisoning for Qwen2.5-14B on OpenMathInstruct and for Llama-3.1-8B on AceReason. Data filtering also proves largely ineffective in detecting poisoned samples: even the best-performing method achieves only 0.378 in F-1 score, leaving the majority of poisoned samples undetected. These findings demonstrate that seemingly benign crowdsourced contributions can amplify leakage of other records while remaining difficult to identify through data filtering.
Figures & tables
| Qwen2.5-7B-It | Llama-3.1-8B-It | Qwen2.5-14B-It | Mistral-Nemo-12B-It | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Metric | Base | SFT | MP | Ours | Base | SFT | MP | Ours | Base | SFT | MP | Ours | Base | SFT | MP | Ours |
| OpenMathInstruct | ||||||||||||||||
| Near-verbatim ER | 0.51 | 1.84 | 3.89 | 6.48 | 0.36 | 2.31 | 0.53 | 2.77 | 0.04 | 2.38 | 2.30 | 8.44 | 0.46 | 1.85 | 2.22 | 4.47 |
| Semantic ER | 1.28 | 2.73 | 4.19 | 7.25 | 0.24 | 4.15 | 1.03 | 5.26 | 0.45 | 3.68 | 2.78 | 10.08 | 0.56 | 2.68 | 4.21 | 6.58 |
| BLEU | 31.0 | 43.7 | 66.6 | 70.8 | 34.0 | 56.1 | 40.7 | 59.0 | 35.4 | 44.2 | 59.5 | 77.8 | 32.5 | 46.6 | 50.3 | 64.0 |
| BLEU | 24.1 | 32.8 | 31.3 | 41.1 | 23.0 | 39.1 | 26.4 | 41.0 | 26.1 | 31.4 | 29.4 | 49.9 | 24.5 | 33.0 | 31.3 | 44.1 |
| Qwen2.5-7B-It | Llama-3.1-8B-It | Qwen2.5-14B-It | Mistral-Nemo-12B-It | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Metric | SFT | Direct | Hosted | SFT | Direct | Hosted | SFT | Direct | Hosted | SFT | Direct | Hosted |
| OpenMathInstruct | ||||||||||||
| Near-verbatim ER | 1.84 | 6.48 | 2.96 | 2.31 | 2.77 | 2.50 | 2.38 | 8.44 | 3.49 | 1.85 | 4.47 | 2.75 |
| Semantic ER | 2.73 | 7.25 | 5.75 | 4.15 | 5.26 | 4.60 | 3.68 | 10.08 | 5.18 | 2.68 | 6.58 | 5.08 |
| BLEU | 43.7 | 70.8 | 53.6 | 56.1 | 59.0 | 58.3 | 44.2 | 77.8 | 52.3 | 46.6 | 64.0 | 61.4 |
| BLEU | 32.8 | 41.1 | 36.0 | 39.1 | 39.8 | 42.0 | 31.4 | 49.9 | 37.2 | 33.0 | 44.1 | 44.1 |
| Qwen2.5-7B-It/OpenMathInstruct | Llama-3.1-8B-It/AceReason | |||||||
| Metric | Base | SFT-Dist. | Poison-Dist. | Proprietary | Base | SFT-Dist. | Poison-Dist. | Proprietary |
| Perplexity ( ) | 1.74 | 1.48 | 1.47 | 1.40 | 1.85 | 1.67 | 1.41 | 1.50 |
| Mean BLEU ( ) | 24.0 | 30.3 | 31.0 | 32.5 | 9.86 | 27.9 | 28.5 | 28.5 |
| Mean Cosine ( ) | 91.1 | 91.6 | 92.1 | 91.9 | 80.7 | 88.7 | 88.5 | 88.5 |
| CVDD ( Ruff et al., 2019 ) | DATE ( Manolache et al., 2021 ) | SIK ( Cao et al., 2025 ) | LLM Judge | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Metric | Inst. | Res. | Inst.+Res. | Inst. | Res. | Inst.+Res. | Inst. | Res. | Inst.+Res. | Inst.+Res. |
| F1 | 0.000 | 0.378 | 0.139 | 0.038 | 0.000 | 0.000 | 0.000 | 0.020 | 0.000 | 0.024 |
| Precision | 0.000 | 0.369 | 0.136 | 0.020 | 0.000 | 0.000 | 0.000 | 0.020 | 0.000 | 0.017 |
| Recall | 0.000 | 0.388 | 0.143 | 0.400 | 0.000 | 0.000 | 0.000 | 0.020 | 0.000 | 0.041 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Topic | Subtopic |
| Math | ||
| Math | Algebra | Linear equations, systems of equations, inequalities, quadratic equations, polynomials, factoring, rational expressions, radicals, exponents, algebraic simplification |
| Math | Number theory | Divisibility, prime numbers, factorization, gcd, lcm, modular arithmetic, congruences, diophantine equations, remainders, parity, digits, floor and ceiling |
| Math | Geometry | Triangles, circles, angles, polygons, area, perimeter, volume, surface area, similarity, congruence, Pythagorean theorem |
| Math | Discrete math | Logic, sets, relations, functions, proof by induction, recurrence relations, graphs, trees, boolean algebra |
| Math | Game theory | Nash equilibrium, zero sum games, dominant strategies, extensive form games, mechanism design, pareto efficiency, backward induction |
| Hyperparameter | Value |
|---|---|
| Learning rate | |
| Epochs | 3 |
| LR scheduler | Cosine, 3% warmup |
| Train batch size (per device) | 1 |
| Gradient accumulation steps | 8 |
| Weight decay | 0.01 |
| OpenMathInstruct | AceReason | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-7B-It | Llama-3.1-8B-It | Qwen2.5-7B-It | Llama-3.1-8B-It | |||||||||||||
| SFT | Poison | SFT | Poison | SFT | Poison | SFT | Poison | |||||||||
| Metric | Prefix | Topic | Prefix | Topic | Prefix | Topic | Prefix | Topic | Prefix | Topic | Prefix | Topic | Prefix | Topic | Prefix | Topic |
| Near-verbatim ER | 0.71 | 1.84 | 5.24 | 6.48 | 1.26 | 2.31 | 2.41 | 2.77 | 0.30 | 1.19 | 0.37 | 1.28 | 0.73 | 1.90 | 1.64 | 6.07 |
| Semantic ER | 1.56 | 2.73 | 6.15 | 7.25 | 2.71 | 4.15 | 3.34 | 5.26 | 0.51 | 2.33 | 2.98 | 4.33 | 3.14 | 3.79 | 6.07 | 10.1 |
| Max BLEU | 46.4 | 43.7 | 68.7 | 70.8 | 55.4 | 56.1 | 58.3 | 59.0 | 28.5 | 30.5 | 29.2 | 41.7 | 37.3 | 52.0 | 36.9 | 62.9 |
| Metric | ||||||
|---|---|---|---|---|---|---|
| Near-verbatim ER | 1.15 | 1.25 | 1.36 | 1.275 | 1.25 | 1.25 |
| Semantic ER | 3.08 | 2.63 | 3.08 | 3.16 | 2.95 | 2.43 |
| Max BLEU | 45.0 | 46.1 | 47.5 | 46.6 | 45.0 | 45.1 |
| Top-10 BLEU | 30.8 | 30.9 | 31.7 | 31.9 | 30.1 | 30.3 |
| OpenMathInstruct | AceReason | ||||
| Domain | Topic | Subtopic | Domain | Topic | Subtopic |
| Top-ranked anchors | |||||
| Math | Algebra | Quadratic equations | Math | Intermediate algebra | Polynomials |
| Math | Number theory | Floor and ceiling | Math | Number theory | Diophantine equations |
| Math | Algebra | Factoring | Math | Trigonometry | Trigonometric equations |
| Math | Trigonometry | Angle additional formulas | Coding | Bit manipulation | Bit counting |
| Target model | |
| Surrogate model | |
| All anchors | |
| Active anchors | |
| Extraction query | |
| Gnerated samples for | |
| Total budget |
| OpenMathInstruct | AceReason | |||||||||||||||
| Qwen2.5-7B-It | Llama-3.1-8B-It | Qwen2.5-7B-It | Llama-3.1-8B-It | |||||||||||||
| Metric | Stat. | Adap. | Stat. | Adap. | Stat. | Adap. | Stat. | Adap. | Stat. | Adap. | Stat. | Adap. | Stat. | Adap. | Stat. | Adap. |
| Near-verbatim ER | 3.62 | 3.99 | 5.93 | 6.48 | 1.88 | 2.06 | 2.75 | 2.77 | 0.77 | 0.79 | 1.34 | 1.28 | 4.47 | 4.84 | 5.99 | 6.07 |
| Semantic ER | 4.39 | 4.82 | 6.90 | 7.25 | 3.26 | 3.48 | 5.16 | 5.26 | 2.53 | 2.73 | 4.29 | 4.33 | 6.66 | 6.76 | 9.86 | 10.1 |
| Original training instruction | Extracted instruction |
|---|---|
| Near-verbatim match (Levenshtein = 0.86, Cosine = 0.99) | |
| You are given an array of strings products and a string searchWord . Design a system that suggests at most three product names from products after each character of searchWord is typed. Suggested products should have common prefix with searchWord . If there are more than three products with a common prefix return the three lexicographically minimums products. Return a list of lists of the suggested products after each character of searchWord is typed . Example 1: Input: products = ["mobile","mouse","moneypot","monitor","mousepad"] , searchWord = "mouse" Output: [["mobile","moneypot","monitor"],["mobile","moneypot","monitor"],["mouse","mousepad"],["mouse","mousepad"],["mouse","mousepad"]] Explanation: Products sorted lexicographically = ["mobile","moneypot","monitor","mouse","mousepad"] . • After typing "m" and "mo" , all products match and we show the top 3: ["mobile", "moneypot", "monitor"] . • After typing "mou" , "mous" , and "mouse" , the system suggests: ["mouse", "mousepad"] . Example 2: Input: products = ["havana"] , searchWord = "havana" Output: [["havana"], ["havana"], ["havana"], ["havana"], ["havana"], ["havana"]] Explanation: The only word "havana" will always be suggested while typing the search word. Constraints: 1 <= products.length <= 1000 1 <= products[i].length <= 3000 1 <= sum(products[i].length) <= 2 * 10^4 All the strings of products are unique . products[i] consists of lowercase English letters. 1 <= searchWord.length <= 1000 searchWord consists of lowercase English letters. Write Python code to solve the problem. Please place the solution code in the following format: # Your solution code here | You are given an array of strings products and a string searchWord . Design a system that suggests at most three product names from products after each character of searchWord is typed. Suggested products should have common prefix with searchWord . If there are more than three products with a common prefix return the three lexicographically minimums products. Return a list of lists of the suggested products after each character of searchWord is typed . Example 1: Input: products = [["mobile","mouse","moneypot","monitor","mousepad"]] , searchWord = "mouse" Output: [["mobile","moneypot","monitor"],["mobile","moneypot","monitor"],["mouse","mousepad"],["mouse","mousepad"],["mouse","mousepad"]] Explanation: Products sorted lexicographically = ["mobile","moneypot","monitor","mouse","mousepad"] . • After typing "m" and "mo" , all products match and we show the top 3: ["mobile", "moneypot", "monitor"] . • After typing "mou" , "mous" , and "mouse" , the system suggests: ["mouse", "mousepad"] . Example 2: Input: products = [["havana"]] , searchWord = "haiti" Output: [] Constraints: 1 <= products.length <= 1000 1 <= products[i].length <= 30 All the strings of products are unique . products[i] consists of lowercase English letters. 1 <= searchWord.length <= 1000 searchWord consists of lowercase English letters. Write Python code to solve the problem. Please place the solution code in the following format: # Your solution code here |
| Semantic match (Levenshtein = 0.66, Cosine = 0.88) | |
| Write a python function to determine if a given string can be segmented into a space-separated sequence of one or more dictionary words. For example, given the dictionary = ["dog", "cat", "code", "sand", "dogcar"] , and the string = "dogcarcode" , the function should return True since the string can be segmented into "dog car code" . However, if the string was "sandydog" , it should return False since no valid segmentation exists. Solve the problem starting with the provided function header. Function header: def word_break(wordDict, s): Please place the solution code in the following format: # Your solution code here | Write a function to check if a given string can be segmented into a space-separated sequence of one or more dictionary words. For example: • Given dictionary = ["dog", "cat"] , sentence = "dogcat" . Return true because it can be segmented as "dog cat" . • Given dictionary = ["dog", "cat"] , sentence = "toad" . Return false because it does not match any dictionary word. Solve the problem starting with the provided function header. Function header: def wordBreak(self, s: str, wordDict: List[str]) -> bool: Please place the solution code in the following format: # Your solution code here |
| Qwen2.5-7B-It | Llama-3.1-8B-It | Qwen2.5-14B-It | Mistral-Nemo-12B-It | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Metric | Base | SFT | MP | Ours | Base | SFT | MP | Ours | Base | SFT | MP | Ours | Base | SFT | MP | Ours |
| Near-verbatim ER | 0.44 | 2.22 | 1.04 | 1.84 | 0.08 | 2.16 | 2.34 | 3.69 | 0.76 | 2.04 | 1.16 | 1.24 | 0.68 | 3.06 | 5.74 | 6.65 |
| Semantic ER | 1.42 | 6.60 | 3.60 | 5.22 | 0.50 | 7.22 | 7.80 | 10.82 | 2.16 | 6.78 | 4.80 | 4.27 | 2.02 | 7.56 | 14.62 | 15.08 |
| Max BLEU | 41.9 | 67.6 | 54.7 | 68.4 | 41.0 | 62.6 | 63.5 | 68.2 | 46.5 | 62.8 | 57.4 | 54.6 | 56.0 | 69.7 | 70.9 | 75.8 |
| Top-10 BLEU | 30.2 | 48.4 | 36.1 | 48.1 | 27.3 | 44.7 | 46.8 | 49.2 | 34.4 | 46.5 | 40.8 | 39.6 | 40.3 | 50.4 | 53.2 | 56.0 |
| Qwen2.5-7B-It | Llama-3.1-8B-It | Qwen2.5-14B-It | Mistral-Nemo-12B-It | |||||||||||||
| Metric | SFT | 50 | 100 | 200 | SFT | 50 | 100 | 200 | SFT | 50 | 100 | 200 | SFT | 50 | 100 | 200 |
| OpenMathInstruct | ||||||||||||||||
| Near-verbatim ER | 1.84 | 4.55 | 6.48 | 8.12 | 2.31 | 2.79 | 3.18 | 3.47 | 2.38 | 8.84 | 8.44 | 11.3 | 1.85 | 3.15 | 4.47 | 5.21 |
| Semantic ER | 2.73 | 6.27 | 7.25 | 8.64 | 4.15 | 5.39 | 5.36 | 5.58 | 3.68 | 9.68 | 10.08 | 12.17 | 2.68 | 5.91 | 6.58 | 7.52 |
| BLEU*max | 43.7 | 63.2 | 70.8 | 79.3 | 56.1 | 59.0 | 57.6 | 63.4 | 44.2 | 78.1 | 77.8 | 89.4 | 46.6 | 59.8 | 64.0 | 71.7 |
| BLEU*10 | 32.8 | 36.5 | 41.1 | 48.1 | 39.1 | 40.8 | 39.8 | 43.1 | 31.4 | 50.6 | 49.9 | 61.5 | 33.0 | 41.8 | 44.1 | 47.7 |
| Qwen2.5-7B / OpenMathInstruct | Llama-3.1-8B / AceReason | |||||||||||
| SFT | Poison | SFT | Poison | |||||||||
| Metric | ||||||||||||
| Near-verbatim ER | 1.12 | 1.84 | 0.90 | 6.79 | 6.48 | 7.51 | 1.80 | 1.90 | 0.56 | 5.96 | 6.07 | 6.34 |
| Semantic ER | 4.16 | 2.73 | 1.68 | 7.54 | 7.25 | 7.46 | 3.24 | 3.79 | 2.26 | 9.38 | 10.1 | 9.43 |
| Max BLEU | 34.9 | 43.7 | 41.1 | 62.2 | 70.8 | 70.6 | 41.8 | 52.0 | 40.1 | 56.0 | 62.9 | 62.3 |
| Top-10 BLEU | 27.3 | 32.8 | 30.9 | 34.2 | 41.1 | 41.5 | 29.1 | 34.6 | 26.9 | 38.8 | 42.7 | 43.0 |
| Qwen2.5-7B-It | Llama-3.1-8B-It | Qwen2.5-14B-It | Mistral-Nemo-12B-It | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Metric | SFT | Direct | Hosted | SFT | Direct | Hosted | SFT | Direct | Hosted | SFT | Direct | Hosted |
| OpenMathInstruct | ||||||||||||
| 1.40 | +0.00 | +0.00 | 1.46 | +0.00 | +0.00 | 1.42 | -0.01 | -0.01 | 1.44 | +0.00 | +0.00 | |
| Mean BLEU ( ) | 35.4 | -2.9 | -0.0 | 32.9 | -0.9 | -0.4 | 32.1 | -2.3 | -2.7 | 34.6 | +0.1 | -0.2 |
| Mean Cosine ( ) | 91.8 | +0.1 | +0.1 | 91.6 | -0.1 | -0.1 | 91.3 | -1.2 | -0.8 | 91.9 | -0.0 | +0.1 |
| AceReason | ||||||||||||
| OpenMathInstruct | AceReason | |||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-7B-It | Llama-3.1-8B-It | Qwen2.5-7B-It | Llama-3.1-8B-It | |||||||||||||||||
| Base | All | Resp. | Base | All | Resp. | Base | All | Resp. | Base | All | Resp. | |||||||||
| Metric | SFT | Poison | SFT | Poison | SFT | Poison | SFT | Poison | SFT | Poison | SFT | Poison | SFT | Poison | SFT | Poison | ||||
| Near-verbatim ER | 0.51 | 1.84 | 6.48 | 1.19 | 1.70 | 0.36 | 2.31 | 2.77 | 0.93 | 1.07 | 0.34 | 1.19 | 1.28 | 0.61 | 0.38 | 0.00 | 1.90 | 6.07 | 0.61 | 0.77 |
| Semantic ER | 1.28 | 2.73 | 7.25 | 1.80 | 2.94 | 0.24 | 4.15 | 5.26 | 1.21 | 1.62 | 0.55 | 2.33 | 4.33 | 1.03 | 1.28 | 0.00 | 3.79 | 10.1 | 1.88 | 2.89 |
| Max BLEU | 31.0 | 43.7 | 70.8 | 37.5 | 45.0 | 34.0 | 56.1 | 59.0 | 44.2 | 48.7 | 22.0 | 30.5 | 41.7 | 23.6 | 27.3 | 20.3 | 52.0 | 62.9 | 22.8 | 41.7 |