cs.LGOct 1, 2026

SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples

Authors: Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik, Aron Laszka

Organizations: College of Information Sciences and Technology, Pennsylvania State University, University Park, PA, USA · Department of Computer Science and Engineering, University of Louisville, Louisville, KY, USA · Department of Computer Science and Engineering, Washington University in St. Louis, St. Louis, MO, USA

Abstract

As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Empirical Evaluation of Data Poisoning Attacks in Supervised Learning

    Sep 10, 2026Toshif Khan, Muhammad AbusaqerPoisoningTraining Data

  2. Are Targeted Data Poisoning Attacks as Effective as We Think?

    Sep 8, 2025William Xu, Chenyu Zhang, Yihan Wang +5PoisoningHard

  3. Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget

    May 2, 2026Yi Yang, Jinyang Huang, Binbin Liu +5Clean Label Backdoor AttackBlack-Box Adversarial Attacks