Organizations: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Imperial College London
Diffusion models have greatly advanced instruction-guided image editing, while also raising concerns about unauthorized image manipulation. Image immunization addresses this risk by adding imperceptible perturbations to an input image to disrupt subsequent edits. Since editing requests are unknown at image release, protection should remain effective beyond the instruction used to construct the perturbation. Existing immunization methods either require costly full-trajectory backpropagation or use intermediate objectives whose effects may be weakened by subsequent denoising. Meanwhile, a single inference path provides limited feedback about alternative denoising continuations. To address these challenges, we propose \textsc{SPIN}, a framework for image immunization via one-step projection over local stochastic trajectory neighborhoods. Starting from an early denoising state, \textsc{SPIN} generates stochastic neighboring states under the same instruction and predicts their clean latents through one-step projection without full unrolling. We then optimize a bounded input perturbation to maximize the average deviation of these predictions from a clean-edit reference, encouraging the perturbation to disrupt multiple possible editing outcomes. Experiments on two image editors demonstrate substantial gains in protection performance, with \textsc{SPIN} outperforming compared methods across all six metrics under seen instructions and in the more challenging unseen instruction setting.
Figures & tables
Figure 1: Overview of SPIN . From a shared native prefix, local trajectory sampling generates neighboring candidates and one-step projection estimates their clean endpoints. These estimates guide perturbation optimization to drive editing outputs away from the fixed clean-edit reference.
Method
SD3
InstructPix2Pix
PSNR ↓
SSIM ↓
VIFp ↓
FSIM ↓
LPIPS ↑
ISR ↑
PSNR ↓
SSIM ↓
VIFp ↓
FSIM ↓
LPIPS ↑
ISR ↑
PhotoGuard-E
18.96
0.6689
0.2134
0.7890
0.3250
58%
13.93
0.4680
0.1010
0.6919
0.5585
83%
PhotoGuard-D
17.16
0.5668
0.1873
0.7682
0.3946
66%
13.16
0.4042
0.1087
0.7053
0.5094
83%
SDS
17.92
0.5868
0.1878
0.7591
0.3798
64%
14.82
0.4704
0.1368
0.7330
0.4703
86%
ED
19.88
0.6530
0.2419
0.8084
0.3176
60%
16.06
0.5098
0.1538
0.7652
0.4199
77%
AdvDM
13.54
0.4519
0.1059
0.6600
0.5391
84%
15.41
0.4839
0.1473
0.7634
0.4388
85%
Table 1: Comparison with image-immunization baselines under original instructions. Lower PSNR/SSIM/VIFp/FSIM and higher LPIPS/ISR indicate better immunization performance.
Method
SD3
InstructPix2Pix
PSNR ↓
SSIM ↓
VIFp ↓
FSIM ↓
LPIPS ↑
ISR ↑
PSNR ↓
SSIM ↓
VIFp ↓
FSIM ↓
LPIPS ↑
ISR ↑
PhotoGuard-E
19.80
0.6928
0.2397
0.8109
0.3007
67%
14.37
0.4884
0.1101
0.7076
0.5486
82%
PhotoGuard-D
20.15
0.7039
0.2471
0.8146
0.2988
64%
16.76
0.4974
0.1565
0.7855
0.4493
81%
SDS
18.79
0.6202
0.2147
0.7843
0.3549
69%
15.54
0.5068
0.1545
0.7536
0.4463
84%
ED
20.64
0.6789
0.2673
0.8305
0.2954
50%
17.10
0.5482
0.1770
0.7951
0.3833
79%
AdvDM
14.36
0.4775
0.1182
0.6834
0.5216
68%
16.65
0.5176
0.1665
0.7833
0.4179
82%
Table 2: Generalization to five unseen instructions. Perturbations are optimized with the original instruction; Lower PSNR/SSIM/VIFp/FSIM and higher LPIPS/ISR indicate better immunization performance.
Figure 2: Qualitative comparisons under seen editing instructions. The first and second rows show examples from InstructPix2Pix and SD3, respectively.
Method
PSNR ↑
SSIM ↑
VIFp ↑
FSIM ↑
LPIPS ↓
ISR ↑
Time/image (s) ↓
SDS
32.85
0.8439
0.5406
0.9700
0.1584
64%
10.7
PhotoGuard-E
31.71
0.8635
0.5405
0.9801
0.0931
58%
8.4
PhotoGuard-D
33.24
0.8473
0.5412
0.9770
0.1835
66%
311.0
AdvDM
33.69
0.8511
0.5540
0.9758
0.1564
84%
14.9
MIST
33.23
0.8433
0.5496
0.9753
0.1611
56%
25.6
SIFM
34.75
0.8836
0.5542
0.9806
0.1279
82%
19.3
Table 3: Input imperceptibility, immunization success, and time cost on SD3.
Variant
Latent cosine dist. ↑
PSNR ↓
LPIPS ↑
SD3
Random perturbation
0.0623
20.78
0.280
Deterministic one-step projection
0.2782
11.70
0.629
Stochastic one-step projection ( K=1 )
0.2814
11.95
0.653
SPIN ( K=4 )
0.3141
11.20
0.673
InstructPix2Pix
Table 4: Effect of one-step projection and stochastic neighborhood aggregation on protection efficacy. The deterministic SD3 variant uses the native deterministic transition, while stochastic variants use one or four candidates.
Figure 4: Effect of normalized sampling progress. (a) We report the PSNR between a candidate’s one-step projected edit and its fully denoised final edit, between projected edits from different candidates, and between the final edit and a fixed clean reference. Lower PSNR indicates a larger difference. (b) Latent cosine distance and PSNR are both computed with respect to the clean edit; higher latent cosine distance and lower PSNR indicate a larger deviation.
Editor
K
Seen PSNR ↓
Seen LPIPS ↑
Unseen PSNR ↓
Unseen LPIPS ↑
Time/image (s) ↓
SD3
1
11.95
0.6532
14.57
0.5702
10.51
SD3
2
10.99
0.6424
14.01
0.5718
12.65
SD3
4
11.20
0.6730
13.95
0.5704
16.01
SD3
8
10.94
0.6479
13.96
0.5828
27.11
InstructPix2Pix
1
11.24
0.6848
15.51
0.5147
5.61
InstructPix2Pix
2
10.88
0.6916
14.36
0.5316
7.40
Table 5: Effect of the number of candidate trajectories at a fixed editor-specific neighborhood scale.
Figure 5: Effect of the perturbation budget ε : (a) input fidelity; (b) output discrepancy under seen and unseen instructions; (c) qualitative examples. Both (a) and (b) report PSNR and LPIPS, with budgets in units of 1/255 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Unseen-instruction comparisons with SD3 on a car image. Columns show the unprotected input and the five protection methods. The top row shows the inputs; the remaining rows show edits under the five instructions printed above them. Each protected input is reused across all five instructions without re-optimization.
Figure 7: Unseen-instruction comparisons with SD3 on a portrait image. The top row shows the unprotected and protected inputs, and the remaining rows show outputs for five instructions excluded from protection optimization. The protected inputs remain fixed across instructions. The comparisons expose changes in subject appearance and image structure beyond the requested edits.
Figure 8: Unseen-instruction comparisons with InstructPix2Pix (IP2P) on a landscape image. The top row shows the unprotected and protected inputs; the remaining rows show edits under five unseen instructions using the same protected inputs. The instructions cover foliage color, mist, lighting, rainbow insertion, and watercolor style.
Figure 9: Unseen-instruction comparisons with InstructPix2Pix (IP2P) on a belt image. The top row shows the unprotected and protected inputs, followed by outputs for five instructions not used during protection optimization. The protected inputs remain fixed across instructions, which request changes to color, hardware, background, text, and texture.
Modern image editors combine vision-language models (VLMs) with diffusion transformer backbones to modify a single reference image according to instructions without fine-tuning. This capability also enables unauthorized manipulation of publicly released images. Existing inference-time defenses either invalidate edits through conspicuous corruption, thereby exposing the protection, or allow them to proceed with identity or reference content drift, thereby failing to prevent the editing behavior itself. We instead target a stealthy and harmless no-op in which the requested edit is suppressed, the output remains natural and source-preserving without conspicuous artifacts or identity replacement, and harmful semantics requested by malicious instructions are absent. We propose NullEdit, which targets the VLM representation jointly formed from the reference image and instruction before it conditions the downstream DiT backbone. Using normal-edit and no-edit anchors, NullEdit redirects this representation, while cross-prompt gradient averaging transfers protection to held out instructions. Across Step1X-Edit and Qwen-Image-Edit on CelebA-HQ and VGGFace2, NullEdit reduces the EditReward IF score by 0.813 on average relative to the SOTA baseline while preserving subject identity and source content.
Weiyao Huang, Liqin Wang, Ziqi Sheng +1
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Diffusion model advances have enabled powerful text-guided image editing, but also raise ethical and legal risks such as deepfakes and unauthorized use. To prevent these risks, adversarial attack-based image immunization has emerged as a promising defense against AI-driven semantic manipulation. Yet, most existing approaches require image-specific optimization or additional neural networks at inference time, hindering scalability and practicality. In this paper, we propose the first universal adversarial perturbation-based image immunization framework that generates a single, image-agnostic adversarial perturbation specifically designed for diffusion-based editing pipelines. Inspired by UAP used in targeted attacks, our method aims to generate a UAP that induces diffusion models to misinterpret the input image as a specific semantic target. Simultaneously, it suppresses original content to misdirect the model's attention during editing, thereby effectively blocking unauthorized edits by overwriting the image's original semantics via the UAP. Extensive experiments show that our method, as the first universal immunization approach, significantly outperforms several baselines in the UAP setting. Notably, despite the inherent difficulty of universal perturbations, our method achieves competitive or superior performance compared to image-specific methods under a more restricted perturbation budget, while also exhibiting strong black-box transferability across diverse diffusion models.
Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent source-image conditioning can limit how fully an edit is executed and how natural the result appears, especially when the target scene diverges substantially from the input. We introduce DuET (Dual Expert Trajectories), a training-free inference method that temporarily relaxes source-image conditioning by transitioning through a text-to-image phase before returning to edit mode, allowing the denoising trajectory to move toward the target distribution while retaining the structural benefits of image-conditioned editing. Without modifying model weights or increasing sampling cost, DuET consistently improves instruction relevance, semantic fidelity, and perceptual quality across diverse models and benchmarks. In some cases, these gains come with a modest reduction in source-image preservation, revealing a predictable trade-off between source preservation and edit fidelity.
Lidia Troeshestova, Alexander Ustyuzhanin, Sergey Kastryulin