cs.CRSep 24, 2026

TraceGuard: Adaptive Multimodal Poison Filtering through Cross-Feature Rank Agreement

Authors: Haoyang Li, Yaxin Xiao, Linyan Dai, Jiawen Fu, Zi Liang, Jason Xue, Qingqing Ye, Haibo Hu

Organizations: The Hong Kong Polytechnic University · Commonwealth Scientific and Industrial Research Organisation (CSIRO)

Abstract

Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We therefore ask which properties a poison set must preserve for the attack to remain effective. A small poison set must still exert enough collective influence during training to induce the attacker's target behavior. We analyze this influence in terms of how often an attack pattern occurs and how strongly the examples carrying it jointly affect the model. This analysis motivates six corpus-level features that examine cross-modal neighborhoods, recurring text, and changes after text-span erasure without training the victim model. We introduce TraceGuard, an adaptive rank-based filtering method that uses agreement among complementary feature rankings to identify suspicious examples. It refines the selected set through shared patterns and adapts the removal threshold to each corpus without knowing the attack or poison rate. Across 19 attack configurations spanning image-text learning, generative vision-language model fine-tuning, and encoder-transfer tests, TraceGuard removes an average of 98.4% of poisoned examples and 5.4% of clean examples. After training on the filtered corpora, the residual attack metric is at most 1% in 13 configurations. Matched-removal controls and ablations support the contributions of sample selection and adaptive removal. Stress tests also identify detection failures under adaptive attacks and unnecessary removal on poison-free corpora.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models

    Oct 4, 2026Tobias Braun, Jonas Henry Grebe, Emil Sivic +4

  2. SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples

    Oct 1, 2026Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik +1PoisoningTraining Data

  3. Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution

    Jun 24, 2026Yan-Lun Chen, Pin-Yu Chen, Chia-Mu Yu +3PoisoningAgentic Retrieval-Augmented Generation Systems