cs.CROct 4, 2026

TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs

Authors: Ruizhi Xu, Wei Xu, Sibo Zhu

Abstract

Backdoor-defense leaderboards print a clean-accuracy drop and read it as removal cost. Measured on the poisoned victim alone, the drop cannot separate removal from what the defense does to any model, and inherits the victim's start, which for three of BackdoorBench's sixteen attacks is a configuration file: WaNet, BPP and Input-Aware ship a MultiStepLR that never fires, so their victims never anneal and are the least accurate in 30/31 public CIFAR cells at ≤\leq5%. On PreAct-ResNet18, fine-tuning-family defenses return a low start to their own level, so there the published cost is negative, the benchmark's rating clips the "gain" to zero, and 2 of 48 citing defense papers we read rest a no-cost claim on those cells; TSBD and CGD, re-run with their code, "gain" on a never-poisoned model too. A 2×22\times2 editing only that scheduler line isolates the cause, its swapped arms self-registered before they ran: the sign of the fine-tuning family's clean-model cost reverses both ways while its published gain on the annealed victim only shrinks toward zero, 44/44 seeds following the schedule, replicated on BPP, FT-SAM, CIFAR-100 and VGG19-BN and induced in a second toolkit. TARE runs the same defense on a never-poisoned twin of the same recipe, schedule and seed (on BackdoorBench, ≤\leq10 poisoned images, admitted only below 5% attack success); what the twin loses is the tare. On the BadNets grid seven of eight defenses charge the twin (Neural Cleanse only where its detector fires), +0.13 (fine-tuning) to +5.70 points (I-BAU); the eighth, ABL, destroys it. Within an attack the start cancels from rankings, so the tare re-orders nothing there; what poisoning adds beyond it is printed under two estimators and not corrected, its removal share unidentified. We ship the three-key patch, a signed tare column (7 attacks ×\times 8 defenses) and TARE-Z, a twin-free estimator for seed-stable defenses.

Explore similar work

CardsList
  1. Density-aware Sample-specific Attack

    May 27, 2026Qiyuan Wang, Yao Li, Raymond K. W. WongBackdoor AttacksAdversarial Robustness

  2. Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor

    Jul 7, 2026Qi Zhao, Christian WressneggerBackdoor AttacksBackdoor Detection

  3. Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget

    May 2, 2026Yi Yang, Jinyang Huang, Binbin Liu +5Backdoor AttacksData Poisoning Attacks