Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the attack under global classifier guidance. We demonstrate three key findings: 1. A carrier mitigates subject distortion by absorbing a larger share of globally normalized attack updates. 2. A carrier improves cross-model transferability, governed by the strength of target-related features that balance semantic separation and transfer performance. 3. Successful targeted attacks retain the personalized subject as the primary content perceived by humans while successfully misleading the classifier. Our results demonstrate that a visually secondary carrier offers an auxiliary spatial pathway for adversarial changes, enabling strong and transferable attacks while improving subject preservation.
Figures & tables
Figure 1: Qualitative comparison of primary subject preservation with and without a visually secondary carrier. Each column shows one source-target case. No Carrier attacks visibly degrade the subject, whereas the shown carrier-based CIRA outputs preserve the personalized subject more faithfully. The first four carrier cases use a Non-Target Carrier, while the final two use a Target Carrier. Both the No Carrier and Carrier-based settings are evaluated under the same strong attack strength and the same attack steps.
WhiteBox
BlackBox
Subject Preservation
Method
Carrier
Top-1 ↑
Top-5 ↑
Cond. Top-5 ↑
Mean Top-5 ↑
DINO ↑
IoU ↑
No-Carrier
None
93.33
98.00
–
2.17
0.9530
0.9940
NatAdiff
N/A
31.83
45.67
–
19.27
0.9411
0.9649
CRA
Non-Target
96.50
99.50
99.49
7.23
0.9857
0.9934
CRA
Hybrid
98.33
99.67
99.64
15.17
0.9818
0.9898
CRA
Target
98.00
99.67
99.56
34.23
0.9845
0.9924
Table 1: Main attack and subject-preservation results over 600 source-target pairs. Cond. Top-5 reports ASR only on samples for which the target class is absent from the clean Top-5 predictions. For NatADiff adaptation, preservation metrics are computed on the 68.17% of pairs for which the designated subject is detected in both the clean and attacked images (Appendix K ). Hybrid rows indicate the requested construction condition; CIRA and JIA do not consistently realize its target-related attributes (Appendix B.5 ).
Attack
Subject Preservation
Method / Condition
Top-1 ASR ↑
DINO ↑
SAM3 IoU ↑
NatADiff
98.83
0.7461
0.9149
No-Carrier
100.00
0.7928
0.9694
Target Carrier ⋅ CRA
100.00
0.9345
0.9666
Target Carrier ⋅ CIRA
100.00
0.9809
0.9695
Table 2: Primary subject preservation under strong attacks. NatADiff adaptation preservation metrics are computed on the 51.00% of pairs for which the designated subject is detected in both the clean and attacked images (Appendix K ).
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Target ID
Target
Target ID
Target
Target ID
Target
94
hummingbird
113
snail
272
coyote
277
red fox
301
ladybug
331
hare
340
zebra
344
hippopotamus
347
bison
348
ram
354
Arabian camel
386
African elephant
404
airliner
407
ambulance
483
castle
487
cellular telephone
504
coffee mug
555
fire engine
Appendix
Table 3: ImageNet-1K target classes used in our experiments.
Target ID
Target
Carrier ID
Base Carrier
113
snail
390
eel
340
zebra
339
sorrel
386
African elephant
344
hippopotamus
404
airliner
417
balloon
407
ambulance
656
minivan
487
cellular telephone
848
tape player
Appendix
Table 4: Target-to-base-carrier mappings used in the formal experiments.
Method
Carrier
Clean Subject Top-5 (%)
Eligible
Cond. Top-5 ASR ↑
CRA
Non-Target
82.35
589/600
99.49
CRA
Hybrid
76.86
562/600
99.64
CRA
Target
82.35
450/600
99.56
CIRA
Non-Target
89.41
593/600
98.48
CIRA
Hybrid
90.78
595/600
98.82
CIRA
Target
89.22
322/600
99.69
Appendix
Table 5: Clean-image WhiteBox prediction audit and conditional Top-5 ASR. Clean Subject Top-5 reports whether the predefined subject category appears in the clean ResNet-50 Top-5 predictions. Eligible denotes the number of samples, out of 600 per setting, for which the target is absent from the clean Top-5 predictions.
Figure 2: Bounded carrier-construction retries for CIRA. Classifier and Qwen feedback identify the missing coffee-mug carrier, after which the prompt and generation seed are varied. None of the four attempts passes the gate; Attempt 2 is retained as the forced fallback based on its best carrier-category rank.
Figure 3: Qualitative examples under the Hybrid Carrier condition for dog5 targeting ”school bus”. The first two columns show the generated carrier background and the composite input, followed by the outputs of CRA, CIRA, and JIA. Green indicates Top-1 attack success, while red indicates failure.
Figure 4: Qualitative examples under the Hybrid Carrier condition for backpack targeting ”ping-pong ball”. The first two columns show the generated carrier background and the composite input, followed by the outputs of CRA, CIRA, and JIA. Green indicates Top-1 attack success.
Figure 5: Carrier occlusion during CRA compositing. Each row shows the generated carrier background, source image, subject insertion mask, and resulting clean composite. Top: severe occlusion, where the cat substantially overlaps the coffee-mug carrier. Bottom: mild occlusion, where most of the elephant carrier remains visible after insertion of the poop-emoji subject.
BlackBox Models
Overall
Method
ResNet-101
VGG-19
Inception-v3
ConvNeXt-B
Swin-B
Macro Avg.
Cond Avg.
CRA
30.17
10.42
10.33
32.67
20.08
20.73
16.71
JIA
37.08
15.08
24.00
41.58
32.67
30.08
26.15
CIRA
35.08
13.58
21.08
37.75
30.33
27.57
23.15
Appendix
Table 6: BlackBox Top-5 results averaged over the Non-Target and Target Carrier conditions. Cond Avg. reports the conditional Top-5 ASR evaluated only on samples whose clean images do not contain the target class in the corresponding model’s Top-5 predictions. For a consistent cross-route comparison, Hybrid is excluded because its requested target-related attributes are not reliably realized by the inpainting-based routes (Appendix B.5 ).
Condition
Subject Top-5 ↑
Target-Absent Top-5 ↑
Non-Target Carrier
82.19
98.67
Hybrid Carrier
80.12
97.03
Target Carrier
83.99
64.10
Appendix
Table 7: BlackBox prediction audit on clean images, averaged over the three attack routes and five BlackBox classifiers (%).
Figure 6: Qualitative comparison of CRA, JIA, and CIRA under the Target-Carrier condition for teapot → castle. The bottom row shows the generated carrier background and target-class Grad-CAM maps for the three attacked outputs.
Figure 7: Non-Target-Carrier example for cat → ambulance. CRA largely occludes the generated carrier; JIA retains the carrier but fails to reach targeted Top-1; CIRA preserves both regions and succeeds. Green and red indicate targeted Top-1 success and failure, respectively.
Figure 8: Example of the carrier-region whitening intervention for a successful CRA attack on dog → ambulance. The visible carrier is localized, its unoccluded region is selected, and the region is whitened. After intervention, the target probability drops from 11.47% to 0.75% , causing the target class to fall from Top-1 to rank 10.
Condition
Initial Success
Failed
Failure Rate
Target Carrier
575
314
54.6%
Non-Target Carrier
555
283
51.0%
All
1130
597
52.8%
Appendix
Table 9: CRA attack failures after whitening the visible carrier region.
Figure 9: Personalized-subject generation failures of the NatADiff-FLUX adaptation under the main setting. The left panel shows the common candle reference used for personalization; it is not an inversion input. Each row shows matched-noise clean and attacked outputs for a different target. Although adversarial guidance is disabled for clean generation, the outputs are dominated by target-related object characteristics and SAM3 cannot detect the designated candle. Consequently, DINO similarity and mask IoU are unavailable. Green and red borders indicate targeted Top-1 success and failure, respectively.
Figure 10: NatADiff adaptation generation under the strong attack setting. The left panel shows the common personalized teapot reference, while each row presents matched-noise clean and attacked outputs for one target. All attacked images achieve targeted Top-1 success, but the teapot characteristics are severely altered or replaced by target-related content. Because SAM3 detection is missing or incomplete, DINO similarity and mask IoU cannot be computed for these pairs.
Setting
Top-1 ASR ↑
Valid Pairs
DINO ↑
IoU ↑
Main
31.83
409/600
0.9411
0.9649
Strong
98.83
306/600
0.7461
0.9149
Appendix
Table 10: NatADiff adaptation attack strength and subject-generation reliability over 600 source–target pairs. Preservation metrics are computed only on valid pairs with successful subject detection in both images.
Evaluation
Subject-Category Selections
Rate (%)
Annotator 1
5,391 / 5,400
99.83
Annotator 2
5,386 / 5,400
99.74
Annotator 3
5,396 / 5,400
99.93
Majority vote
5,398 / 5,400
99.96
Appendix
Table 11: Human-evaluation interface. Annotators are shown an attacked image and asked to select its primary category from four randomly ordered options: the personalized-subject category, the attack target, and two distractor categories. The reported selection rate is the proportion of responses matching the personalized-subject category.
Figure 11: Screenshot of the human-evaluation interface. Annotators are shown an attacked image and asked to identify its primary subject from four randomly ordered candidate categories.
Transfer-based adversarial attacks rely on surrogate models to craft perturbations, yet often overfit the surrogate's decision boundary. To address this problem, we propose Inverse Knowledge Distillation (IKD), a simple and attack-agnostic mechanism that maximizes the prediction-distribution discrepancy between benign and adversarial samples on the surrogate model. IKD uses a CE/KL-equivalent soft-label objective to push adversarial predictions away from a fixed benign prediction anchor and enrich the attack with Fisher-sensitive surrogate directions. We prove that, under a matched fixed-anchor implementation, soft-label cross-entropy and KL divergence differ only by a constant entropy term and therefore induce identical gradients, Hessians, and adversarial optimization trajectories. Our information-geometric analysis further derives a quantitative lower bound on dominant Fisher-subspace overlap between surrogate and target models from local same-task stability and a Fisher eigengap, and establishes a sufficient target-margin crossing condition under oriented gradient coherence and target smoothness. This analysis connects IKD's surrogate Fisher sensitivity to cross-model transfer. In contrast, mean squared error uses a different Euclidean pullback in output probability space. IKD integrates seamlessly with standard gradient-based attacks without modifying their optimization pipelines. Extensive ImageNet experiments demonstrate consistent black-box gains across CNN, ViT, and defended models, while ablations confirm CE and KL equivalence and the pronounced disadvantage of MSE. These results establish IKD as an effective and lightweight component for improving adversarial transferability. Code is available at https://github.com/ImmortalTing/IKD.
Wenyuan Wu, Yuan Sun, Yingke Chen +4
College of Computer Science, Sichuan University, China · Department of Computer and Information Sciences, Northumbria University, UK · School of Artificial Intelligence, Sichuan University, China
Transfer-based adversarial attacks often transfer poorly across heterogeneous architectures because CNNs favor local textures while Vision Transformers (ViTs) rely on global shapes. We propose Season, a spectrum-aware orthogonal gradient refinement framework for L-infinity transfer attacks against black-box target models on ImageNet, using a white-box surrogate. Season decomposes each update into a low-frequency branch capturing structural cues and a high-frequency branch capturing textures. A low-saliency guidance scheme reallocates high-frequency energy to background regions, preserving foreground structures that ViTs depend on. An orthogonal projection then forces the textural update to lie in the orthogonal complement of the structural direction, mitigating feature interference. As a training-free plug-and-play wrapper, Season enhances eight gradient-stabilization and input-enhancement attacks without modifying their cores. Across eight CNN, ViT, and MLP targets, Season improves transfer success rate by 6.6 percentage points on average and up to 16.0 points over strong baselines under a unified protocol.
Tianyi Wang, Zhenghao Gao, Shengjie Xu
Tongji University Shanghai, China · Huazhong University of Science and Technology · Wuhan LightRead Intelligent Technology Co., Ltd. Wuhan, China
Adversarial examples reveal vulnerabilities in Vision-Language Pre-training (VLP) models and provide insights for improving robustness. A key property is cross-model transferability, which enables transfer-based black-box attacks. However, existing attacks often rely heavily on the surrogate model, causing cross-model performance drops. One reason is that adversarial optimization may follow surrogate model responses more than input semantics, making the update direction effective on the surrogate but less transferable to unseen targets. We refer to this dependency as surrogate-specific bias. Motivated by this observation, DeBias-Attack improves transferability by correcting surrogate-specific bias in adversarial optimization directions. It maintains two perturbation branches. The main branch optimizes a perturbation on the original image and obtains the adversarial gradient used to disrupt image-text alignment. The reference branch optimizes a perturbation on a weak-semantic image constructed from the dataset mean image with small Gaussian noise resampled at each iteration. Since this weak-semantic image contains little clear visual content, its optimization reflects surrogate responses more than image semantics, and its reference gradient estimates surrogate-specific bias. DeBias-Attack removes the aligned projection of the main gradient on the reference gradient before updating the adversarial image, then performs context-aware text substitution using the updated adversarial image. DeBias-Attack is the first transfer-based VLP attack that corrects surrogate-specific bias through gradient correction. Experiments show strong performance across VLP models, downstream tasks, and open-source and closed-source multimodal large language models.
Lijia Yu, Jiuxin Cao, Yuchen Qiang +3
School of Cyber Science and Engineering, Southeast University, China · Purple Mountain Laboratories, Nanjing, China