Invisible watermarking has become a central tool for tracing AI-generated images, but its robustness against adaptive removal attacks remains an open security question. We introduce Latent Frequency Masking, an attack that erases watermark evidence by replacing selected Fourier coefficients in the latent representation of a watermarked image. The replacement can be sampled from Gaussian noise for efficiency or derived from diffusion regeneration for improved image preservation. We provide a theoretical distortion bound relating the change between the reconstructed adversarial image and the masked latent-frequency perturbation. We evaluate the proposed attack against six diffusion watermarking methods on images generated from DiffusionDB and MS-COCO prompts. Latent Frequency Masking removes or substantially weakens several watermarks while preserving perceptual quality and achieving favorable runtime compared with existing attacks. These results identify latent-frequency manipulation as a practical attack surface and highlight the need to include such attacks in robustness evaluations of generative image watermarking.
Figures & tables
Fig. 1: Overview of the watermark robustness evaluation setup
Fig. 2: The overall scheme of the proposed Latent Frequency Masking removal algorithm
Fig. 3: Removal efficiency (TPR@1%FPR ↓ of the watermark after attack) vs image quality (PSNR ↑ ) for different attacks measured on DiffusionDB dataset. Lower-right is better; LFM variants sit on the Pareto front for all methods except Gaussian Shading
DiffusionDB
MS-COCO
Robustness
Quality metrics
Robustness
Quality metrics
Attack
TPR@1%FPR ↓
PSNR ↑
SSIM ↑
LPIPS ↓
CLIP IQA ↑
TPR@1%FPR ↓
PSNR ↑
SSIM ↑
LPIPS ↓
CLIP IQA ↑
No attack
0.97
∞
1.000
0.000
0.866
0.98
∞
1.000
0.000
0.883
Adv. Embedding [ 21 ]
0.71
22.2
0.462
0.402
0.349
0.71
22.3
0.459
0.439
0.301
Diff. Regeneration [ 23 ]
0.51
21.7
0.620
0.205
0.840
0.51
22.4
0.657
0.178
0.833
Diff. Rinsing [ 21 ]
0.42
20.1
0.553
0.258
0.856
0.44
20.8
0.591
0.223
0.838
TABLE I: Attack performance and quality metrics averaged across all tested watermarks on DiffusionDB and MS-COCO datasets
Fig. 4: Removal efficiency (TPR@1%FPR ↓ of the watermark after attack) vs image quality (LPIPS ↑ ) for different attacks measured on DiffusionDB dataset. Lower-left is better; LFM variants sit on the Pareto front for all methods except Gaussian Shading
Fig. 5: Average attack performance vs execution time. The size of the dot represents the average PSNR on the DiffusionDB dataset
Fig. 6: Examples of a Tree-Ring-watermarked image subjected to different attacks. The prompt is sampled from the DiffusionDB database.
Attack
TPR@1%FPR
PSNR
VAE
1.0
33.1
VAE + Masking
1.0
23.7
FFT + Masking
0.995
11.4
VAE + FFT + Masking (Gaussian)
0.001
19.4
VAE + FFT + Masking (Diffusion)
0.385
28.7
TABLE II: Ablation results, VAE represents only the Encoder and Decoder steps without any masking, VAE + Masking: masking performed in spatial domain without FFT step, FFT + Masking: frequency masking is performed on images without Encoder and Decoder steps, VAE + FFT + Masking: full algorithm
Digital image watermarking has advanced rapidly for copyright protection of generative AI, yet the comparatively limited progress in watermark attack techniques has broken the attack-defense balance and hindered further advances in the field. In this paper, we propose FMDiffWA, a frequency-domain modulated diffusion framework for watermark attacks. Specifically, we introduce a frequency-domain watermark modulation (FWM) module and incorporate it into the sampling stages both the forward and reverse diffusion processes. This mechanism enables selective modulation of watermark-related frequency components, thereby allowing FMDiffWA to effectively neutralize the invisible watermark signals while preserving the perceptual quality of the attacked watermarked images. To achieve a better trade-off between attack efficacy and visual fidelity, we reformulate the training strategy of conventional diffusion models by augmenting the canonical noise estimation objective with an auxiliary refinement constraint. Comprehensive experiments demonstrate that FMDiffWA achieves superior visual fidelity compared to existing watermark attacks, while exhibiting strong generalization across diverse watermarking schemes.
Chunpeng Wang, Binyan Qu, Xiaoyu Wang +4
Qilu University of Technology (Shandong Academy of Sciences) Jinan, Shandong Province, China · Dalian Maritime University Dalian, Liaoning Province, China · Nanjing University of Science and Technology Nanjing, Jiangsu Province, China +1
The proliferation of AI-generated images produced by Latent Diffusion Models (LDMs) has raised critical concerns regarding copyright infringement and misinformation. Although existing frequency-domain watermarking methods embed handcrafted geometric patterns into the initial latent noise prior to generation, they suffer from limited capacity and rigid pattern designs. We propose DeepFreqMark, an end-to-end learnable frequency-domain watermarking framework that replaces manual pattern engineering with a neural message encoder and decoder. To circumvent the computational bottleneck caused by Denoising Diffusion Implicit Model (DDIM) inversion during training, we introduce a Spherical Linear Interpolation (Slerp)-based attack simulation. This approach operates directly on the noise latent while strictly preserving the Gaussian variance. Extensive experiments demonstrate that DeepFreqMark achieves significantly lower Bit Error Rates (BER) than baseline methods under real-world attacks and scales to 256 bits message capacity. Our source code is available at https://github.com/chenhsiu48/DeepFreqMark.
Chen-Hsiu Huang, Mario Köppen, Ja-Ling Wu
National Taiwan University, Taipei, Taiwan · Kyushu Institute of Technology, Fukuoka, Japan
Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored. Existing attacks either succeed only against specific generative models or achieve removal at the cost of severe visual degradation. In this paper, we propose MarkNull, a model-agnostic watermark removal attack via on-manifold latent manipulation. MarkNull is grounded in a key observation: watermarked images exhibit a strong statistical dependency between the generated latent representation and the embedded initial noise. To quantify this dependency, we introduce the Noise-Latent Alignment Score (NLAS) and formulate an optimization objective that selectively decorrelates the latent representation from the embedded watermark while preserving semantic fidelity. Extensive evaluations across different categories of watermarking paradigms, including post-hoc, fine-tuning-based, and initial-noise-based schemes, demonstrate that MarkNull reduces average bit accuracy to 53.14%, approaching random-guessing (50%), without perceptible image degradation. To further improve scalability, we propose MarkNull-A, an amortized, optimization-free variant that distills the attack into a single forward pass, achieving 0.50 s/image with modest computational overhead. Notably, our attacks successfully compromise Google's SynthID-Image system while preserving high visual quality and transfer effectively to video watermarking. Finally, we present an attack detection mechanism as a defensive counterpart to MarkNull and MarkNull-A, highlighting the necessity of developing watermark designs resilient to model-agnostic latent-space attacks.
Jie Cao, Qi Li, Zelin Zhang +4
Queen’s University, Canada · University of Waterloo, Canada