Watermarking diffusion-generated images requires balancing provenance signals with image quality, robustness, and computational cost. This work organizes methods along two axes: insertion mechanism and primary signal-bearing representation, and formalizes a representative zT-Fourier pipeline for verification and identification. We then use the taxonomy to structure three protocol-bounded case studies. The first examines associations among frequency integrity, detection, quality, and cropping behavior. The second revisits persistence under seed-linked and seed-independent editing and formulates a scoped Semantic Imprinting Hypothesis without claiming a localized carrier or causal mechanism. The third studies single-shot VAE-latent phase modulation, including its efficiency, regeneration robustness, and robustness--quality operating points. Finally, we separate four content-level attack families from model/pipeline adaptation, propose corresponding evaluation protocols and testable conjectures for parameter-tuning threats, and identify additional temporal extensions for video. These analyses do not establish a universal ranking; instead, they provide a framework for matched, protocol-aware comparisons of watermarking systems for diffusion-generated images.
Table 1 : Taxonomy matrix of generative-image watermarking methods by insertion mechanism and primary signal-bearing representation ( Sec. 2.3 ). Reproduced from [ 16 ] .
Figure 1 : Representative generation-time zT -Fourier watermarking pipeline used in the case studies: a keyed Fourier pattern is inserted into the initial noise, propagated through diffusion generation, recovered from an optionally transformed query via DDIM inversion, and compared with reference keys for verification or closed-set identification. Reproduced from [ 16 ] , which is itself adapted from [ 15 ] . The diagram does not represent post-hoc optimization, spatial- zT , z0 , or pixel-domain methods, and the optional transformation is illustrative rather than a robustness claim.
Method
Post
Payload / IDs
Verif.
Ident.
Δ FID-1k
DwtDctSvd [ 5 ]
✓
32/106
0.494
0.206
−0.389
RivaGAN [ 28 ]
✓
32/106
0.692
0.484
−0.785
Stable Signature [ 7 ]
×
48/106
0.819
0.489
−0.241
Tree-Ring [ 25 ]
×
zero / 2048
0.699
0.125
+1.184
HSTR [ 15 ]
×
zero / 2048
0.993
0.946
−0.099
RingID [ 4 ]
×
zero / 2048
0.999
0.977
+2.054
Table 2 : Detection over the harmonized attack subset and generative quality under the PhaseMark configuration. Except for HSTR, values follow [ 17 ] ; HSTR detection is recomputed from [ 15 ] without brightness and noise, and its Δ FID-1k is our matched-condition measurement. “Post” denotes post-hoc applicability, and “Payload/IDs” reports payload bits and enrolled-pool size. Verification is TPR@1%FPR. Identification follows each source-native rule: nearest-pattern accuracy for zero-bit methods or thresholded bit-decoding TPR@1%FPR for payload methods. Lower Δ FID-1k is better. Adapted from [ 16 ] .
Method
0.6
0.5
0.4
0.3
RingID [ 4 ]
0.971
0.919
0.774
0.559
HSTR [ 15 ]
1.000
0.999
0.992
0.903
HSQR [ 15 ]
1.000
1.000
0.999
0.999
Table 3 : Closed-set identification accuracy under random-crop attacks at four crop scales; smaller scales indicate more aggressive cropping [ 15 ] . Adapted from [ 16 ] .
Category
Method
P2P
InfEdit
Spectral Modification
Tree-Ring [ 25 ]
0.25
0.32
Optimization-based
ZoDiac [ 29 ]
0.75
0.86
Frequency-aligned
HSQR [ 15 ]
0.99
1.00
Cryptographic
G.Shading [ 27 ]
1.00
0.98
Table 4 : Verification TPR@1%FPR after Prompt-to-Prompt (P2P) and InfEdit editing [ 18 ] . Watermarked source images are generated with SD2.1, whereas InfEdit uses SD1.5, yielding a cross-model editing condition. This editing-specific protocol differs from the harmonized general-attack protocol in Tab. 2 , so absolute values are not directly comparable across the two tables. Category labels are study-specific descriptors rather than classes in Tab. 1 . Reproduced from [ 16 ] .
Figure 2 : Schematic outcomes used to illustrate SIH. The upper example has low pre-edit recoverability and is not detected after editing, while the lower example has high pre-edit recoverability and is detected after editing. The figure does not identify the watermark carrier or show that recoverability before editing determines the detection result. Adapted from [ 16 ] ; the underlying study is [ 18 ] .
Method
Space
Embedding
Detection
Regen. avg.
ZoDiac [ 29 ]
zT
∼7.3 min (opt.)
4.11 s (inv.)
0.958
G.Shading [ 27 ]
zT
Sampling
4.11 s (inv.)
0.999
HSQR [ 15 ]
zT
Sampling
4.11 s (inv.)
0.997
PhaseMark-APM [ 17 ]
z0
0.14 s (one-shot)
0.05 s
0.998
Table 5 : Embedding procedure or per-image cost, detection latency, and mean verification TPR@1%FPR across VAE-B, VAE-C, and diffusion regeneration [ 17 ] . “Sampling” denotes generation-integrated insertion; “opt.” and “inv.” denote per-image optimization and DDIM inversion. Reproduced from [ 16 ] .
Modulation
Strategy Type
Verif.
Ident.
PSNR ↑
LPIPS ↓
(T@1%F)
(T@1%F)
APM (Robust)
Hard / Absolute
0.999
0.988
31.715
0.129
IPS
Hard / Relative
0.998
0.980
32.437
0.090
SPS
Soft / Relative
0.996
0.960
32.952
0.068
PCQ (Quality)
Soft / Absolute
0.970
0.816
34.156
0.037
Table 6 : PhaseMark variants: mean verification and identification TPR@1%FPR over clean data and nine attacks, with image-quality metrics [ 17 ] . Higher PSNR and lower LPIPS are better. Reproduced from [ 16 ] .
Attack family
Modified object
Evidence in this paper
Content-level attacks Fixed embedding rule, key, and detector
Signal processing
Samples of an existing image
General attacks [ 15 , 17 ]
Cropping
Spatial support of an existing image
Crop analysis [ 15 ]
Generative editing
Source-conditioned re-synthesis
P2P; cross-model InfEdit [ 18 ]
Regeneration
Re-synthesis by VAE or LDM
Regeneration study [ 30 , 22 , 17 ]
Table 7 : Operational threat families and evidence scope. Content-level attacks transform an existing watermarked image under a fixed embedding rule, key, and detector; model/pipeline adaptation changes future watermark behavior and is supported here only by method-specific reports. Adapted from [ 16 ] .
Invisible watermarking has become a central tool for tracing AI-generated images, but its robustness against adaptive removal attacks remains an open security question. We introduce Latent Frequency Masking, an attack that erases watermark evidence by replacing selected Fourier coefficients in the latent representation of a watermarked image. The replacement can be sampled from Gaussian noise for efficiency or derived from diffusion regeneration for improved image preservation. We provide a theoretical distortion bound relating the change between the reconstructed adversarial image and the masked latent-frequency perturbation. We evaluate the proposed attack against six diffusion watermarking methods on images generated from DiffusionDB and MS-COCO prompts. Latent Frequency Masking removes or substantially weakens several watermarks while preserving perceptual quality and achieving favorable runtime compared with existing attacks. These results identify latent-frequency manipulation as a practical attack surface and highlight the need to include such attacks in robustness evaluations of generative image watermarking.
Kirill Aistov, Khaled Abud, Irina Serzhenko +8
MSU AI Institute, Moscow, Russia · Trusted AI Research Center RAS, Moscow, Russia
Digital image watermarking has advanced rapidly for copyright protection of generative AI, yet the comparatively limited progress in watermark attack techniques has broken the attack-defense balance and hindered further advances in the field. In this paper, we propose FMDiffWA, a frequency-domain modulated diffusion framework for watermark attacks. Specifically, we introduce a frequency-domain watermark modulation (FWM) module and incorporate it into the sampling stages both the forward and reverse diffusion processes. This mechanism enables selective modulation of watermark-related frequency components, thereby allowing FMDiffWA to effectively neutralize the invisible watermark signals while preserving the perceptual quality of the attacked watermarked images. To achieve a better trade-off between attack efficacy and visual fidelity, we reformulate the training strategy of conventional diffusion models by augmenting the canonical noise estimation objective with an auxiliary refinement constraint. Comprehensive experiments demonstrate that FMDiffWA achieves superior visual fidelity compared to existing watermark attacks, while exhibiting strong generalization across diverse watermarking schemes.
Chunpeng Wang, Binyan Qu, Xiaoyu Wang +4
Qilu University of Technology (Shandong Academy of Sciences) Jinan, Shandong Province, China · Dalian Maritime University Dalian, Liaoning Province, China · Nanjing University of Science and Technology Nanjing, Jiangsu Province, China +1
Diffusion watermarking embeds verifiable signals into the generative process and commonly verifies them by recovering trajectory-dependent evidence, making the marks robust to conventional pixel-space distortions. Existing removal attacks either regenerate along deterministic trajectories, which often preserve the watermark-bearing latent structure, or optimize every image separately. We identify the reliance on a recoverable generative trajectory as a common attack surface among the schemes we study. Based on this observation, we propose DRIFT, a black-box attack that combines partial forward diffusion with stochastic reverse resampling. Forward re-noising limits source information available to a fixed-depth recovery pipeline, while stochastic reversal supplies alternative noise-driven paths whose removal benefit we isolate through matched sampler comparisons. Adaptive DRIFT searches a selected ladder for each image's first verifier-rejected rung and refines fidelity while retaining only updates rejected by the same verifier. At fixed depth, we derive information-theoretic and Wasserstein source-dependence bounds; under realized-ladder monotonicity, the first rejected rung is least distorted among rejected rungs on that ladder, and verifier-gated refinement preserves rejection. Across nine watermarks spanning three paradigms, DRIFT achieves 98-100% attack success and the best image quality among the compared attacks, without secret keys, verifier internals, or per-image gradient optimization.
Rui Bao, Zheng Gao, Xiaoyu Li +3
University of New South Wales · Griffith University