Organizations: College of Information Sciences and Technology, Pennsylvania State University, University Park, PA, USA · Department of Computer Science and Engineering, University of Louisville, Louisville, KY, USA · Department of Computer Science and Engineering, Washington University in St. Louis, St. Louis, MO, USA
As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.
Figures & tables
Attack
No Defense
SAGE (Ours)
Deep k -NN
Meta-Sift
EPIC
Confusion Training
Size
Method
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
CIFAR-10
ϵ = 8
FC
0.9460 ± 0.0008
15%
0.9258 ± 0.0017
0%
0.9389 ± 0.0034
0%
0.9233 ± 0.0003
15%
0.9465 ± 0.0008
0%
0.9474 ± 0.0042
5%
BP
0.9457 ± 0.0007
100%
0.9366 ± 0.0012
0%
0.9471 ± 0.0002
20%
0.9155 ± 0.0002
25%
0.9466 ± 0.0008
10%
0.9416 ± 0.0046
45%
GM
0.9487 ± 0.0006
75%
0.8980 ± 0.0025
0%
0.8973 ± 0.0123
10%
0.8907 ± 0.0009
20%
0.8787 ± 0.0048
10%
0.9040 ± 0.0065
10%
ϵ = 16
FC
0.9456 ± 0.0008
25%
0.9392 ± 0.0024
0%
0.9426 ± 0.0032
10%
0.9255 ± 0.0003
25%
0.9465 ± 0.0006
0%
0.9451 ± 0.0032
5%
TABLE I: Post-defense attack success rate (ASR) and clean test accuracy (Acc.) for clean-label triggerless poisoning attacks on CIFAR-10 and Tiny ImageNet. Results are averaged over 20 poisoned datasets for each attack and perturbation budget ϵ . ASR is the fraction of successful attacks across the 20 poisoned datasets, while Acc. is reported as mean ± standard deviation. For each attack, the best result among the defenses is shown in bold, and the second-best result is underlined.
Attack
No Defense
SAGE (Ours)
Activation Clustering
SPECTRE
Confusion Training
FLARE
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
Acc. ( ↑ )
ASR ( ↓ )
CIFAR-10
Narcissus
0.9521 ± 0.0004
0.9120 ± 0.0010
0.9281 ± 0.0030
0.1310 ± 0.0017
0.9023 ± 0.0020
0.0431 ± 0.0025
0.9045 ± 0.0006
0.1235 ± 0.0290
0.9149 ± 0.0010
0.0987 ± 0.0061
0.5364 ± 0.0294
0.0452 ± 0.0011
Wicked-BadNets
0.9467 ± 0.0006
0.8970 ± 0.0035
0.9494 ± 0.0022
0.2720 ± 0.0048
0.9469 ± 0.0006
0.8915 ± 0.0067
0.9425 ± 0.0055
0.8959 ± 0.0174
0.9475 ± 0.0049
0.8065 ± 0.0074
0.5507 ± 0.0392
0.0699 ± 0.0469
Wicked-Blended
0.9471 ± 0.0002
0.9383 ± 0.0007
0.9467 ± 0.0007
0.4620 ± 0.0030
0.9447 ± 0.0029
0.9381 ± 0.0014
0.9428 ± 0.0047
0.5608 ± 0.0758
0.9408 ± 0.0042
0.5675 ± 0.0054
0.4586 ± 0.0040
0.1023 ± 0.0762
Wicked-SIG
0.9451 ± 0.0007
0.9372 ± 0.0018
0.9401 ± 0.0022
0.3400 ± 0.0157
0.9407 ± 0.0056
0.9382 ± 0.0262
0.9400 ± 0.0054
0.7613 ± 0.0475
0.9452 ± 0.0041
0.7612 ± 0.0038
0.4941 ± 0.0354
0.0717 ± 0.0206
TABLE II: Post-defense clean test accuracy (Acc.) and attack success rate (ASR) for four clean-label backdoor attacks on CIFAR-10. Results are mean ± std over 10 runs. For each attack, the best result among the defenses is shown in bold, and the second-best result is underlined. Note that FLARE’s low ASR comes at the cost of excessively low clean accuracy.
Fig. 1: Effect of verified-set composition on poison recall and false positives. The upper panel reports the number of poisoned images identified, and the lower panel reports the number of clean images removed. Solid lines denote class-balanced sampling of verified clean examples. Dashed lines denote unstratified random sampling.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Scale τ
Loss
Accuracy
64
13.5630
2.3252
0.7020
128
13.0431
2.2413
0.7146
256
13.0456
2.0417
0.7328
512
13.0370
2.1955
0.7242
Appendix
TABLE IV: Effect of the feature-extractor output dimension d . Cross-entropy loss and accuracy are measured on the clean CIFAR-100 test set using the similarity-weighted prediction rule in Equation ( 3 ).
∣R∣
Scale τ
Accuracy
22
0.0000
0.5166
100
1.2895
0.6386
200
3.5260
0.7350
500
2.0171
0.6867
Appendix
TABLE V: Effect of the reference-set size ∣R∣ used during feature-extractor training. Accuracy is measured on the clean CIFAR-100 test set using the similarity-weighted prediction rule in Equation ( 3 ).
Data poisoning corrupts training data to degrade a model or to plant attacker-controlled behavior. This study evaluates two representative training-time attacks, label flipping and backdoor poisoning, on MNIST and Fashion-MNIST with three baseline classifiers: Logistic Regression, Linear SVM, and Random Forest. Clean training is compared with poisoning rates of 5%, 10%, and 20% using clean-test accuracy, macro-precision, macro-recall, macro-F1, and, for backdoors, attack success rate. Label flipping caused clear degradation, largest for Logistic Regression and Linear SVM, while Random Forest stayed comparatively stable. Backdoor poisoning reached attack success rates from 0.9667 to 1.0000 on both datasets and all three models while often keeping clean-test performance near baseline. The results separate indiscriminate poisoning, which shows up in standard metrics, from targeted backdoor poisoning, which stays comparatively stealthy while embedding highly effective malicious behavior, and they support security-oriented evaluation beyond conventional clean-test metrics.
Toshif Khan, Muhammad Abusaqer
Minot State University · Department of Math, Data & Technology Minot State University Minot, ND, 58707 · Trustworthy Language Intelligence Lab, Minot State University
Targeted data poisoning attacks manipulate model predictions on specific test samples by injecting malicious data into training. Yet existing evaluations report average attack success rates over randomly selected targets, obscuring true worst-case effectiveness. We argue that the right evaluation focuses on the hardest samples to poison. The same reasoning applies to defense: since targeted attacks leave no footprint at the distribution level, defenders should proactively identify the most vulnerable samples and apply targeted countermeasures. Given a test dataset, this paper identifies both the easiest and hardest to poison examples based on only clean model information. Specifically, we offer coarse evaluations using clean training dynamics, and fine-grained classification on poison class using poison distances and budgets. Our experiments show these metrics reliably stratify samples by poisoning vulnerability, enabling both rigorous worst-case evaluation and proactive vulnerability-aware defense.
William Xu, Chenyu Zhang, Yihan Wang +5
Waabi AI · University of Waterloo · Carnegie Mellon University +3
Backdoor attacks threaten the deep learning supply chain by poisoning a small fraction of the training data so that a model behaves normally on clean inputs but misclassifies trigger-carrying inputs to an attacker-chosen target class. Clean-label backdoor attacks are especially dangerous because poisoned samples remain label-consistent and are therefore harder to detect. Yet existing clean-label attacks typically rely on expensive optimization, surrogate-model training, or nontrivial data access. We present Checkerboard, a theoretically grounded, learning-free clean-label backdoor attack that is effective, efficient, and simple to implement. From a linear separability formulation, we derive a checkerboard trigger in closed form, removing the need for surrogate-model training and trigger optimization. For texture-rich datasets, we introduce Complexity-driven Sample Selection, which uses only target-class data to improve trigger-to-background contrast by selecting low-complexity images for poisoning. Across four benchmark datasets, Checkerboard outperforms 8 baseline attacks and achieves state-of-the-art performance under low poisoning budgets. For example, on CIFAR-10, under a trigger perturbation budget of 10/255, poisoning 20 training samples achieves 99.99% Attack Success Rate (ASR). On ImageNet-100, a poisoning rate of only 0.46% yields over 94% ASR without degrading clean accuracy. The proposed attack also remains effective against state-of-the-art backdoor defenses and shows strong resistance to adaptive defenses.
Yi Yang, Jinyang Huang, Binbin Liu +5
School of Computer Science and Information Engineering Hefei University of Technology Hefei, Anhui, China · School of Information Science and Technology University of Science and Technology of China Hefei, Anhui, China · Faculty of Business Data Science Kansai University Japan +3