WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks
Authors: Khaled Abud, Aleksey Yakushev, Aleksandr Akimenkov, Irina Serzhenko, Kirill Aistov, Egor Kovalev, Dmitry Obydenkov, Sergey Lavrushkin, +4 more
Organizations: MSU Institute for Artificial Intelligence Moscow, Russia · Trusted AI Research Center RAS Moscow, Russia · Independent researcher Moscow, Russia
Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated content and ensuring traceable sources to prevent manipulation or misuse. Recent advances in invisible watermarking methods highlight the need to update existing benchmarking practices to reflect current techniques and evaluation criteria. We address this by introducing WARP -- a unified framework and benchmark for evaluating the robustness of invisible watermarks. WARP incorporates 32 recent classical, deep, and generative watermarking methods, as well as 34 different erasing techniques, ranging from traditional distortions to more sophisticated adversarial, purification, and re-embedding attacks. It provides standardized, reproducible, and easily scalable protocols for evaluating perceptual quality, watermark readability, and attack resilience. Using WARP, we extensively evaluate current invisible watermarking techniques, collecting the largest robustness benchmark in the field. Results identify the most robust approaches under both distortion and adversarial conditions, and reveal consistent relationships between watermarking methods and the attack strategies most effective against them. Our experiments also highlight that some of the watermarking methods considered are highly vulnerable to reembedding, even if they are robust to standard distortions. The code is made available at https://github.com/ispras/wibe.
Figures & tables
Figure 1. Visualizations of multiple invisible watermarking methods available in our benchmark. The second row shows the pixel-wise difference from the source image (magnified ×7 for visibility) alongside the embedded message capacity in bits. A grid of image thumbnails. The top row shows a source photograph of an archer silhouetted against a sunset, followed by its watermarked versions produced by nine methods: HiDDeN, StegaStamp, DWSF, DFT Circle, ARWGAN, DWT DCT, MBRS, ChunkySeal and InvisMark. The bottom row shows the corresponding pixel-wise differences from the source, magnified seven times, with each method's message capacity printed underneath, ranging from zero-bit to 1024 bits. All watermarked images are visually indistinguishable from the source, while the difference maps reveal method-specific signatures: dense high-amplitude texture for HiDDeN and StegaStamp, block and ring structures for the frequency-domain methods, and almost entirely black maps for ChunkySeal and InvisMark.
Name
WA Methods
Attacks
Metrics
Benchmark
Post-hoc
Built-in
Total
Recent methods
Stirmark ( Kutter and Petitcolas, 1999 )
0
0
0
0
20
4
invisible-watermark ( ShieldMnt, 2021 )
3
0
3
0
0
1
✗
SSL watermarking ( Fernandez et al., 2022 )
2
0
2
0
6
4
✗
WAVES ( An et al., 2024 )
3∗
2∗
5
0
20
9
✓
W-Bench ( Lu et al., 2025 )
11
0
11
2
13
7
✓
Table 1. Comparison of frameworks for watermarking research. "Benchmark" column indicates whether the framework includes comprehensive results for the implemented methods, with denoting limited evaluations. * next to the number denotes that the count is based on the methods reported in the original paper, but the number of fully implemented watermarking pipelines found in public repository is smaller. We consider methods recent if they have been released after Jan 2025.
Figure 2. Overview of the WARP framework pipeline, inherited from WIBE ( Yakushev et al., 2025 ) . The process begins with (1) Watermark Embedding and (2) Quality & Imperceptibility Evaluation on the left. The central module, (3) Attack Simulation, applies diverse perturbations or adversarial transformations. On the right, the pipeline continues with (4) Watermark Extraction, (5) Post-Attack Evaluation, and (6) Aggregation & Logging, while (7) Visualization & Reporting is performed as the final step. A block diagram of the WARP evaluation pipeline, flowing from left to right. On the left, source images together with an embedded bit-string enter either a post-hoc invisible watermarking algorithm, of which 26 are available, or a diffusion model with built-in watermarking driven by text prompts, of which 6 are available; both produce watermarked images, and the difference from the source feeds a quality and imperceptibility evaluation. In the centre, an attack module applies one of the available erasure techniques, grouped into traditional transformations such as crop, JPEG, blur, noise, resize and rotate; image re-generation such as VAE, Deep Image Prior and Stable Diffusion; adversarial attacks such as watermark inversion and SADRE; and purification attacks such as DiffPure, DISCO, NRP and RealESRGAN. The result is images with corrupted watermarks. On the right, a watermark decoder extracts a bit-string that is compared with the embedded one, and results are logged and passed to a data analysis and visualisation stage covering text-image, single-image and multi-image quality metrics and watermark robustness metrics including BER, WER and true and false positive rates.
Figure 3. Capacity-quality tradeoff for all tested post-hoc watermarking methods on clear images from MS-COCO. Color of the point represents execution speed, and marker determines the watermark type. Blue line indicates the Pareto-optimal front. * indicates algorithms that only use CPU. Scatter plot of embedding capacity in bits on a logarithmic x-axis from 1 to 1024 against LPIPS distortion on the y-axis from 0 to 0.05, where lower is better, for all tested post-hoc watermarking methods on MS-COCO. Marker shape encodes the architecture family and marker colour encodes combined embedding and extraction time. StegaStamp sits alone at the top with by far the highest distortion. Earlier methods such as SSL, HiDDeN, DWT SVM and DWT DCT occupy a low-capacity region with visible distortion, while recent methods including PixelSeal, VideoSeal, MBRS, InvisMark and ChunkySeal form a Pareto front along the bottom right, combining capacities of 100 to 1024 bits with LPIPS below 0.01. Colour also separates the two computational regimes, with the classical CPU-only methods slower than the GPU-accelerated learned ones.
Figure 4. (a) Robustness-distortion tradeoff for different post-hoc watermarking methods averaged across all tested attacks on DiffusionDB. Robustness is measured with TPR@0.01%FPR ( ↑ ); (b) Per-attack robustness breakdown for all tested watermarks. Two panels. Panel (a) is a scatter plot of average true positive rate at 0.01 percent false positive rate across all attacks on the x-axis against PSNR in decibels on the y-axis, for post-hoc watermarking methods on DiffusionDB, with marker colour encoding embedding capacity. Classical and early methods including DWT DCT, DWT SVM and HiDDeN cluster on the left at low robustness; StegaStamp reaches high robustness but at the lowest PSNR of about 30 decibels; and Robust-Wide, PixelSeal, MaskWM and InvisMark occupy the upper-right frontier, combining high imperceptibility with high robustness. Panel (b) is a heatmap whose rows are all tested watermarking methods, split into a generative block and a post-hoc block, and whose columns are the evaluated attack configurations grouped into traditional distortions, adversarial purification, erasure and regeneration. Cell colour gives the true positive rate at 0.01 percent false positive rate from 0 to 1. Most methods retain high values under traditional distortions and purification, whereas the regeneration columns, the non-right-angle rotation column and SADRE are almost uniformly dark, indicating near-complete watermark loss. The generative watermarks, in particular Gaussian Shading and MaXsive, remain bright across the regeneration columns.
Figure 5. Watermark robustness under re-embedding attacks . (a) Average TPR@0.01%FPR ( ↑ ) measured for each pair of victim and attacker watermarks. (b) Bit Error Rate ( ↓ ) of multi-bit watermarks under re-embedding attacks, averaged across all attacker watermarks. (c) Watermark efficiency as an attacker is evaluated as the average BER induced across watermarks. Three panels on re-embedding attacks. Panel (a) is a matrix heatmap whose rows are victim watermarks, with post-hoc methods above and generative methods in a separate block below, and whose columns are attacker watermarks; cell colour gives the average true positive rate at 0.01 percent false positive rate. Most cells are bright, showing broad resistance to external re-embedding, but the diagonal is dark for nearly every post-hoc method, showing that re-applying the same method erases the original mark; key-based methods and DCT CAISS are the exceptions, and the generative rows stay uniformly bright. Panel (b) is a bar chart of the average bit error rate induced under re-embedding for each victim method, sorted in increasing order from Gaussian Shading near zero to DWT DCT at roughly 0.15. Panel (c) is a bar chart of the average bit error rate each attacker method induces across all victims, sorted in decreasing order and led by HiDDeN, StegaStamp, PIMoG and MaskWM, the methods with the lowest PSNR.
Figure 6. Efficiency-distortion tradeoff for different attacks averaged across all watermarks on DiffusionDB dataset. Attack efficiency is measured with BER. Attacks with colored names lie close to the efficiency-visibility Pareto front. Scatter plot of attack effectiveness against image fidelity for all evaluated attacks on DiffusionDB, with average bit error rate across watermarking methods on the x-axis from 0 to about 0.45 and SSIM on the y-axis from 0.2 to 1.0. Marker shape encodes the attack group and colour encodes average runtime; an inset zooms into the crowded low-BER region. Dedicated erasure attacks including WMForger, TrustMark-RM, DISCO and DIP lie along the upper-right Pareto front, combining high bit error rates with minimal image change. Regeneration attacks such as Flux Regeneration, Flux Rinsing and SADRE reach the highest bit error rates but the lowest SSIM values. Geometric transforms, particularly 30 and 90 degree rotation and centre cropping, also combine high bit error rates with low SSIM.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Name
Year
Type
Binary message
Bits
NN-based
Architecture
Frequency Domain
Latent Space
Local/Global
Video
DWT DCT ( Al-Haj, 2007 )
2007
post-hoc
yes
100 bits
no
-
yes (DWT, DCT)
no
local
no
DWT DCT SVD ( Navas et al., 2008 )
2008
post-hoc
yes
100 bits
no
-
yes (DWT, DCT)
no
local
no
DFT Circle ( Poljicak et al., 2011 )
2011
post-hoc
no
zero-bit
no
-
yes (FFT)
no
local
no
DCT CAISS ( Guzik et al., 2015 )
2011
post-hoc
yes
800 bits
no
-
yes (DCT)
no
local
no
DWT SVM ( Islam et al., 2020 )
2020
post-hoc
yes
512 bits
no
Support Vector Machine
yes (DWT)
no
local
no
HiDDeN ( Zhu et al., 2018 )
2018
post-hoc
yes
30 bits
yes
CNN, GAN
no
no
global
no
Appendix
Table 2. Watermarking methods and their properties. Type denotes whether the watermarking algorithm processes an existing image or is built into the process of image generation. Local/global denotes whether watermark affects only a portion of / all pixels (if it is added in pixel space) or coefficients of the domain it is inserted in (e.g. Fourier coefficients or latent space). Video denotes whether the algorithm was designed to support video in addition to images.
Figure 7. Theoretical vs. empirical TPR@FPR for multibit watermarking under Gaussian noise. Solid lines represent empirical TPR estimated using a reference set of N=100,000 non-watermarked images from MS-COCO train split following the WAVES protocol; Dashed lines represent theoretical TPR computed from the binomial model assuming i.i.d. Bernoulli(0.5) bits. Four FPR thresholds are shown: 10% , 1% , 0.1% and 0.01% . Results are reported for TrustMark (k=100 bits) on the MS-COCO validation split. Line plot of TPR on the y-axis from 0 to 1 against Gaussian noise standard deviation on the x-axis from 0.10 to 0.40, for TrustMark with a 100-bit payload on the MS-COCO validation split. Four pairs of curves correspond to false positive rate thresholds of 10 percent, 1 percent, 0.1 percent and 0.01 percent; solid lines give the empirical TPR estimated from a reference set of 100,000 non-watermarked images and dashed lines the theoretical value from the binomial model. All curves decrease monotonically as noise increases, and within each pair the theoretical curve lies slightly above the empirical one, with the gap widening at higher noise levels and at looser thresholds.
Figure 8. TPR@0.01%FPR vs. Bit Error Rate (BER). Each point corresponds to a multibit watermarking method. The x-axis reports BER (lower is better), the y-axis reports TPR at a fixed FPR of 0.01% (higher is better). Marker colors denote different payload capacities. Results are aggregated across all evaluated attacks on the DiffusionDB dataset. Scatter plot with average bit error rate across all attacks on the x-axis from 0.10 to 0.35, where lower is better, and true positive rate at a fixed 0.01 percent false positive rate on the y-axis, with one point per multi-bit watermarking method and colour encoding payload capacity. The overall relationship is negative, but at comparable bit error rates the high-capacity methods, including ChunkySeal with 1024 bits and PixelSeal with 256 bits, achieve substantially higher true positive rates than 30-bit methods such as ARWGAN and Riva GAN. DWT DCT and HiDDeN sit in the lower-right corner with the worst values on both axes.
FID
CLIP-IQA
Aesthetic
ImageReward
CLIPScore
BLIP
Marked
Non-Marked
Δ↓
Val. ↑
Rel. Δ , % ↑
Val. ↑
Rel. Δ , % ↑
Val. ↑
Rel. Δ , % ↑
Val. ↑
Rel. Δ , % ↑
Val. ↑
Rel. Δ , % ↑
Watermark
Stable Signature
23.29
23.49
-0.20
0.86
-4.24
5.34
-0.45
0.45
0.30
0.27
0.58
0.54
0.25
Gaussian Shading
25.17
25.27
-0.09
0.92
-0.10
5.27
0.05
0.42
-0.20
0.26
-0.12
0.53
-0.22
RingID
25.19
25.43
-0.24
0.92
-0.20
5.27
-0.16
0.42
-0.02
0.26
-0.06
0.53
-0.23
MaxSive
25.41
25.27
0.14
0.92
0.25
5.25
-0.52
0.42
-0.09
0.27
0.12
0.53
-0.04
Appendix
Table 3. Image generation quality evaluation for built-in watermarking methods. For each method, images are generated both with and without the watermark using the same underlying generative model. FID is computed against COCO validation images. Val. column represents the absolute value of the corresponding metric, and Rel. Δ indicates the change relative to the base generative model. See Section C.2 for more details on Δ scores calculation.
Figure 9. TPR@xFPR heatmaps for multibit watermarking methods. Two heatmaps show TPR@xFPR across decreasing FPR thresholds from 10−1 to 10−8 . Left: method-wise heatmap. Rows are multibit watermarking methods, sorted by decreasing average TPR across all attacks and FPR levels. Columns are FPR thresholds. Each cell shows TPR at the corresponding FPR. Right: attack-wise heatmap. Rows are attack types, sorted by decreasing average TPR across all methods and FPR levels. Columns represent FPR thresholds. Each cell shows TPR averaged over all multibit watermarking methods for that attack at the given FPR. Results are reported on the DiffusionDB dataset. Two heatmaps sharing a column axis of eight false positive rate thresholds decreasing from 10 percent to 0.000001 percent. In the left heatmap, rows are multi-bit watermarking methods sorted by decreasing average true positive rate, with Gaussian Shading, StegaStamp, PixelSeal and VideoSeal at the top and SSL, HiDDeN and DWT DCT at the bottom. Values fall as the threshold tightens, sharply for low-capacity methods such as HiDDeN, which drops from about 0.8 to near 0, and only mildly for high-capacity methods such as ChunkySeal and PixelSeal. In the right heatmap, rows are attacks sorted by decreasing average true positive rate across methods; LIIF, deblurring and averaging attacks at the top leave the true positive rate close to 1, whereas rotation, SADRE, Flux Regeneration and Flux Rinsing at the bottom drive it close to 0 at every threshold.
Figure 10. Full robustness breakdown measured by Average BER ( ↓ ) for all evaluated watermarking methods across all 40 attack configurations on the DiffusionDB dataset. Each cell reports the mean BER over 1k images; values close to 0.5 indicate complete watermark erasure. Watermarks are sorted by average BER, and attacks are arranged by their average adversarial effectiveness. Heatmap of average bit error rate on the DiffusionDB dataset. Rows are all evaluated watermarking methods, with generative watermarks in a separate block at the top and post-hoc methods sorted by increasing average bit error rate; columns are the attack configurations, grouped into traditional distortions, adversarial purification, erasure and regeneration and ordered by increasing adversarial effectiveness. Cells near 0 are pale and cells near 0.5, corresponding to complete watermark erasure, are dark. Most cells in the traditional distortion and purification groups are close to 0, while the regeneration columns and the rotation, SADRE and WMForger columns approach 0.5 for nearly all post-hoc methods. The DWT DCT, DWT SVM and HiDDeN rows are the darkest overall, and the Gaussian Shading and Robust-Wide rows the palest.
Figure 11. Full robustness breakdown measured by Average BER ( ↓ ) for all evaluated watermarking methods across all 40 attack configurations on the MS-COCO dataset. Each cell reports the mean BER over 1k images; values close to 0.5 indicate complete watermark erasure. Heatmap of average bit error rate on the MS-COCO dataset, with the same layout as the DiffusionDB breakdown: rows are post-hoc watermarking methods sorted by increasing average bit error rate and columns are attack configurations grouped into traditional distortions, adversarial purification, erasure and regeneration. The pattern closely matches the DiffusionDB results, with near-zero values under traditional distortions and purification for most methods, and values approaching 0.5 under regeneration, non-right-angle rotation and dedicated erasure attacks. Robust-Wide and PIMoG have the lowest average bit error rate, and DWT SVM and DWT DCT the highest.
Figure 12. Full robustness breakdown measured by Average WER ( ↓ ) for all evaluated watermarking methods across all 40 attack configurations on the MS-COCO dataset. Each cell reports the mean WER over 1k images; values close to 1.0 indicate complete watermark erasure. Heatmap of average word error rate on the MS-COCO dataset, where a value of 1.0 means every message is lost. Rows are post-hoc watermarking methods sorted by increasing average word error rate and columns are attack configurations grouped into traditional distortions, adversarial purification, erasure and regeneration. Because a single flipped bit invalidates the whole message, the heatmap is far more saturated than the corresponding bit error rate figure: entire columns for regeneration attacks, rotation and several erasure attacks are at 1.0 for nearly every method, and the HiDDeN, HiDDeN (SS) and DWT SVM rows are close to 1.0 almost everywhere.
Figure 13. Robustness scores of all tested multi-bit watermarks across samples from 5 datasets: DiffusionDB, MS-COCO, NIPS2017, KonIQ and DIV2K. Image resolution varies from 299x299 in NIPS2017 up to 2040x1300 in DIV2K. The scores are averaged across all attacks on each dataset. Three radar charts, one for each robustness metric: average bit error rate, average true positive rate at 0.01 percent false positive rate, and average word error rate. In each chart, every axis corresponds to one of the 24 multi-bit post-hoc watermarking methods and each coloured polygon to one of five datasets: NIPS2017 at 299 by 299 pixels, DiffusionDB at 512 by 512, MS-COCO at about 640 by 480, KonIQ-10k at 1024 by 768, and DIV2K at 2040 by 1300. The polygons are nearly concentric, showing that the relative ranking of methods is essentially the same on all five datasets, while the higher-resolution datasets sit consistently at slightly better values. The largest spread between datasets occurs for DWT SVM and HiDDeN (SS), and the word error rate chart is visibly more compressed than the other two.
Figure 14. Quality–robustness tradeoff for post-hoc watermarking methods evaluated under four imperceptibility metrics: PSNR (dB, ↑ ), SSIM ( ↑ ), LPIPS ( ↓ ), and CLIP-IQA ( ↑ ). Robustness is reported as average BER across all attacks ( ↓ ). Color encodes embedding capacity in bits. While the overall rankings are broadly consistent across metrics, notable divergences between full-reference (PSNR, SSIM, LPIPS) and no-reference (CLIP-IQA) measures reveal that different metrics capture complementary aspects of watermark visibility. Four scatter panels sharing an x-axis of average bit error rate across all attacks, where lower is better, and plotting on the y-axis, respectively, PSNR in decibels, SSIM, LPIPS and CLIP-IQA, for post-hoc watermarking methods, with colour encoding embedding capacity in bits. Rankings are broadly consistent across the four metrics: methods with low bit error rates also tend to score well on perceptual quality. InvisMark attains the highest PSNR of any method at about 50 decibels, and StegaStamp is worst on PSNR, SSIM and LPIPS. Two divergences stand out: DWT SVM has one of the worst LPIPS values despite an unremarkable SSIM, and HiDDeN (SS) has the lowest CLIP-IQA of any method at about 0.73 despite mid-range PSNR and SSIM.
Figure 15. Perceptual quality–adversarial effectiveness tradeoff for all evaluated attacks, measured under three quality criteria: LPIPS ( ↓ ), CLIP-IQA ( ↑ ), and Aesthetic score ( ↑ ). Attack effectiveness is reported as average BER across all watermarking methods ( ↑ ). Color encodes average attack runtime. Regeneration-based attacks (crosses) achieve high BER with strong no-reference quality scores, indicating that they replace rather than degrade the image. Geometric distortions incur high LPIPS but mostly preserve semantic quality. Three scatter panels sharing an x-axis of average bit error rate across all watermarking methods, where higher means a more effective attack, and plotting LPIPS, CLIP-IQA and aesthetic score on the y-axes for every evaluated attack. Marker shape encodes attack group and colour encodes average runtime. Regeneration attacks, in particular Flux Regeneration and Flux Rinsing, reach the highest bit error rates with elevated LPIPS but also the highest CLIP-IQA and aesthetic scores, indicating that they re-synthesise rather than degrade the image. Rotations by 30 and 90 degrees show high LPIPS while leaving the no-reference scores largely intact. GaussBlur with deblurring and DogBlur score poorly on both effectiveness and quality, while WMForger and TrustMark-RM lie on the favourable frontier of all three panels.
Figure 16. BER–SSIM tradeoff curves averaged across watermarks for each evaluated attack under hyperparameter variation on the COCO dataset. Higher SSIM indicates less image distortion; higher BER indicates more effective watermark removal. Line plot of average bit error rate on the y-axis from 0 to about 0.45 against SSIM on the x-axis from 0.1 to 1.0, averaged over all watermarking methods on the MS-COCO dataset, with one curve per attack whose main hyperparameter is swept: two Adversarial Embedding variants, Gaussian blur, JPEG compression, Gaussian noise, random crop-out and WMForger. Every curve rises as SSIM falls, so the most efficient attacks lie towards the upper right. WMForger dominates, reaching high bit error rates at SSIM above 0.8, followed by JPEG compression and the two Adversarial Embedding variants, while Gaussian blur, Gaussian noise and random crop-out need far more image distortion to reach comparable bit error rates.
Figure 17. BER–distortion curves under hyperparameter variation for six attacks, shown per watermarking method on the COCO dataset. Each panel varies the attack’s primary hyperparameter: learning rate (WMForger), Linf budget ϵ (Adversarial Embedding), quality factor (JPEG), std (Gaussian noise), and kernel size (Gaussian blur). Distortion is measured by LPIPS ( ↓ ). Six panels in two rows of three, each plotting bit error rate on the y-axis from 0 to 0.5 against LPIPS distortion on the x-axis, with one curve per watermarking method, as the primary hyperparameter of an attack is varied on the MS-COCO dataset. The top row shows erasure attacks: WMForger with varying learning rate and Adversarial Embedding with CLIP and with ResNet backbones with varying perturbation budget. The bottom row shows traditional distortions: JPEG compression with varying quality factor, Gaussian blur with varying kernel size, and Gaussian noise with varying standard deviation. Under WMForger nearly every method transitions sharply to chance-level bit error rate within a narrow LPIPS window of roughly 0.12 to 0.20, whereas the Adversarial Embedding curves are gradual and widely spread. DCT CAISS stays far below all other methods in most panels and shows a delayed threshold jump under Gaussian blur, while InvisMark degrades unusually quickly under light JPEG compression.
Figure 18. Examples of images generated with different built-in watermarking methods compared to the ones generated by the same underlying generative model with the same prompt and seed without watermark insertion. The prompt used for all presented generations: "A black cat sits in a car and looks out." Six pairs of generated photographs of a black cat looking out of a car window, one pair for each built-in watermarking method: Gaussian Shading, MaXsive, METR, RingID, Stable Signature and TreeRing. In each pair the left image is watermarked and the right image was produced by the same generative model with the same prompt and seed without watermark insertion. Watermarked and non-watermarked images are of comparable visual quality with no visible artefacts. For the noise-space methods the two images of a pair differ in composition and framing rather than in fidelity, whereas Stable Signature, which watermarks through decoder fine-tuning, preserves the composition almost exactly.
Figure 19. Frequency-domain signatures of post-hoc watermarking methods, visualized as the average FFT spectrum difference between watermarked and source images on the COCO subset. Amplitude is shown on a logarithmic scale. A grid of 24 frequency-domain maps, one per post-hoc watermarking method, showing the average FFT amplitude difference between watermarked and source images on the MS-COCO subset, on a logarithmic colour scale spanning four orders of magnitude. Classical methods produce highly structured signatures: a prominent X-shaped cross for DCT CAISS, a regular grid-like lattice for DWT SVM, a bright square frame at the image boundary for DWSF and VideoSeal, and a uniformly bright field for DFT Circle. Among learned methods, StegaStamp shows a strong ring pattern at mid-to-high frequencies and TrustMark a structured square motif, while FIN and CIN show faint fine-grained dot patterns. InvisMark, MaskWM, WAM and PixelSeal are almost featureless, indicating spectrally diffuse embeddings.
Attack
Year
Type
Brief description
Brightness
-
Distortion
Changes image intensity levels, making embedded patterns harder to preserve or detect
Center Crop
-
Distortion
Removes outer regions of an image, potentially discarding embedded information
Color Inversion
-
Distortion
Flips pixel values, altering image patterns and disrupting embedded information
Contrast
-
Distortion
Rescales intensity differences, altering pixel relationships and distorting image patterns
Gaussian Blur
-
Distortion
Smooths the image using a Gaussian filter, reducing detail and weakening fine patterns
JPEG
-
Distortion
Lossy quantization that removes high-frequency details and introduces compression artifacts
Appendix
Table 4. All attacking methods used in our study alongside their brief descriptions.
Invisible image watermarks are increasingly used for deepfake detection and provenance tracking, where they must survive not only incidental distortions but also deliberate removal. We revisit spread-spectrum embedding, a classical watermarking principle, inside a modern neural post-hoc watermarking architecture. Our starting point is a measurement: in existing encoder-decoder schemes each message bit occupies only a small fraction of the image, a shared contributing factor to their fragility, since removal then need only disturb the region a bit occupies. SpreadMark instead spreads each bit as a dense pseudo-random codeword over the whole image and recovers it by matched-filtering a learned cover-suppressed chip representation, with a parallel convolutional decoding path and sparsification-aware training. A conditional chip-space analysis shows that, under a codeword-independent perturbation model, dense spreading increases the budget required to disrupt matched-filter recovery. Evaluated on COCO and DIV2K against nine schemes, SpreadMark is the only evaluated method retaining high detection under both the regeneration and the latent-space sparsification settings we test, with competitive JPEG and additive-noise robustness. It keeps the embedded watermark imperceptible, maintaining high perceptual quality on both COCO and DIV2K.
Wei Song, Yuxin Cao, Zhenchang Xing +4
University of New South Wales, Australia · National University of Singapore, Singapore · CSIRO’s Data61, Australia
Local image watermarking embeds an invisible signal into selected image regions rather than spreading it across the entire image, enabling payload recovery from specific objects or regions without perceptibly altering the image. Existing studies evaluate the robustness of payload recovery and localization under image transformations, but they often focus on their own proposed method, resulting in narrow evaluations with inconsistent choices of transformations, datasets, and metrics. These inconsistencies across studies limit direct comparisons across methods and muddle the overall picture of local watermark robustness. To address this gap, we present the first systematic robustness benchmark for local watermarks across 55 image transformations, including (i) signal distortions, (ii) changes in image coordinate alignment, (iii) indirect local edits, and (iv) direct watermark edits. The benchmark evaluates MaskWM, WAM, OmniGuard, TrustMark, and PixelSeal, all methods that either provide native localization or require minimal adaptation to support it. Our results show that all evaluated methods are vulnerable to some transformation, with MaskWM standing out as offering the strongest payload recovery and localization, although it has the lowest image quality in the clean setting. Synchronization further improves MaskWM's payload recovery under several geometric transformations, albeit at an additional cost to image quality. A key finding is that local watermark robustness depends strongly on the nature of the transformation: signal distortions are often tolerated by the strongest methods, while geometric misalignment and generative local edits, such as inpainting and outpainting, can completely impair payload recovery. We observe that payload recovery and localization are related but not interchangeable, and both strongly depend on the transformation's impact on the watermark region.
Kai Yao, Bence Szilágyi, Sebestyén Kamp +4
University of Edinburgh · Garandor · Garandor; University of Edinburgh
Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored. Existing attacks either succeed only against specific generative models or achieve removal at the cost of severe visual degradation. In this paper, we propose MarkNull, a model-agnostic watermark removal attack via on-manifold latent manipulation. MarkNull is grounded in a key observation: watermarked images exhibit a strong statistical dependency between the generated latent representation and the embedded initial noise. To quantify this dependency, we introduce the Noise-Latent Alignment Score (NLAS) and formulate an optimization objective that selectively decorrelates the latent representation from the embedded watermark while preserving semantic fidelity. Extensive evaluations across different categories of watermarking paradigms, including post-hoc, fine-tuning-based, and initial-noise-based schemes, demonstrate that MarkNull reduces average bit accuracy to 53.14%, approaching random-guessing (50%), without perceptible image degradation. To further improve scalability, we propose MarkNull-A, an amortized, optimization-free variant that distills the attack into a single forward pass, achieving 0.50 s/image with modest computational overhead. Notably, our attacks successfully compromise Google's SynthID-Image system while preserving high visual quality and transfer effectively to video watermarking. Finally, we present an attack detection mechanism as a defensive counterpart to MarkNull and MarkNull-A, highlighting the necessity of developing watermark designs resilient to model-agnostic latent-space attacks.
Jie Cao, Qi Li, Zelin Zhang +4
Queen’s University, Canada · University of Waterloo, Canada