To comply with recent regulations requiring traceable generated content, modern watermarking has adopted multi-bit post-hoc watermarking schemes. These modern designs rest on an encoder-decoder pair implemented as deep neural networks. These models are usually treated as pure black-boxes trained end-to-end, with the noise of the watermarking channel modeled through a fixed set of geometric and valuemetric transforms applied to watermarked images. We argue that this purely empirical approach leads to unquestioned design flaws and a lack of theoretical performance guarantees. This work proposes a general theoretical model of modern post-hoc watermarking schemes grounded in a statistical analysis of the outputs of the encoder/decoder pair. We show that these deep neural networks implicitly define a watermarking channel modeled as parallel AWGN channels, with messages transmitted using BPSK modulation. This imposes a binary alphabet, greatly limiting the capacity of these watermarking systems. Another fatal flaw is their lack of a secret key, making them intrinsically insecure. We make this notion of watermarking security precise for post-hoc schemes by linking it to the possibility of estimating the secret key under a given statistical model of the decoder's output. By putting together the results from this theoretical analysis, we introduce SNW: a novel post-hoc watermarking system that significantly outperforms existing state-of-the-art baselines in terms of capacity while also providing strong security guarantees. Notably, it does not depend on a fixed codebook or binary alphabet, allowing it to reach a rate close to Shannon capacity through the use of capacity-achieving error-correcting codes.
Figures & tables
Figure 1: Empirical analysis of the decoding projection fd and decision mechanism d of three SOTA post-hoc watermarking systems computed over 10k 1024×1024 ImageNet Deng et al. (2009) images watermarked with random messages. (Left) Eigenvalues of the covariance ΣIdentity of the latent vector fd(xwm) (i.e. when t is the identity). (Right) Global distribution of the soft-codeword values c~ and the corresponding theoretical AWGN model with (ϱ,σ) computed empirically.
Channel characteristic ϱσ−1 ( ↑ )
Capacity C(ϱ,σ) / Empirical Capacity
Identity
JPEG 50
Contrast ×2
Crop 0.6
Rot. 90 ∘
Identity
JPEG 50
Contrast ×2
Crop 0.6
Rot. 90 ∘
PixelSeal
3.07
2.21
1.60
1.73
1.97
0.99 / 0.96
0.90 / 0.88
0.69 / 0.72
0.75 / 0.75
0.84 / 0.79
VideoSeal
2.03
1.64
1.03
0.71
1.17
0.85 / 0.85
0.71 / 0.75
0.38 / 0.47
0.21 / 0.21
0.74 / 0.75
TrustMark
3.11
2.67
1.26
0.26
0.20
0.99 / 0.98
0.96 / 0.95
0.52 / 0.58
0.03 / 0.00
0.02 / 0.00
Table 1: Channel characteristic ϱσ−1 computed empirically over 10k 1024×1024 ImageNet images watermarked with random messages and the corresponding theoretical capacity in message bits per codeword bit.
Figure 2: Theoretical bit accuracy p(ρ) (left) and Shannon capacity in bits (right) as a function of the projection cosine alignment ρ . Setting α=M′/L guarantees perfect security against PCA attacks while maintaining an invariant bit accuracy across all codeword dimensions M′ . Dotted lines represent the maximum robustness setting, when α=1 . Security ratio η required to estimate the secret subspace with a PCA attack as a function of α for different values of alignment ρ and M′∈{256,512} .
Figure 4
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Losses
Stage 1
Stage 2
Stage 3
Lalignall
✓
✗
✗
Lalignpool
✗
✓
✓
Lisosame
✓
✓
✓
Lisocross
✗
✓
✓
Lisohost
✗
✓
✓
LqualPSNR
✓
✓
✓
Appendix
Table 3: Active losses per training stage.
WM
Projection function
Lipschitz constant Lf
PixelSeal
ConvNeXT
1079
VideoSeal
ConvNeXT
595
TrustMark
ResNet
120
Broken-Arrows
DWT
1
Appendix
Table 4: Comparison of empirical estimates of the Lipschitz constant Lf of the projection function across schemes. Classical transforms guarantee distance preservation ( Lf=1 ), whereas unconstrained neural projection functions exhibit large empirical Lipschitz constants, exposing them to low-distortion removal. The estimation is performed over 100 ImageNet images.
Method
Rate Rσ / Rσ×M′
Identity
Residual Transfer
PixelSeal
0.910 / 232.9
0.003 / 0.8
VideoSeal
0.848 / 217.2
0.002 / 0.6
Trustmark
0.506 / 50.6
0.059 / 5.9
SNW (w/o anti-spoofing)
0.883 / 678.0
0.113 / 86.5
SNW (w anti-spoofing)
0.888 / 682.2
0.000 / 0.2
Appendix
Table 5: Bit accuracy of the methods against residual transfer attacks. The results are computed over 200 MFlickr 1024×1024 images.
Figure 5: Eigenvalues of the covariance ΣIdentity of the latent vector fd(xwm) (i.e. when t is the identity) for different checkpoint of SNW. The number in the legend correspond to the number of training step.
Method
Identity
VAE
Sana (2 steps)
Sana (4 steps)
Sana (8 steps)
PixelSeal
0.939 / 240.5
0.555 / 142.0
0.519 / 132.8
0.357 / 91.3
0.076 / 19.4
VideoSeal
0.884 / 226.2
0.576 / 147.6
0.557 / 142.6
0.392 / 100.4
0.084 / 21.5
TrustMark
0.632 / 63.2
0.453 / 45.3
0.415 / 41.5
0.184 / 18.4
0.023 / 2.3
SNW
0.892 / 684.7
0.495 / 380.3
0.479 / 368.1
0.388 / 298.3
0.145 / 111.5
Method
Sana (10 steps)
Sana (20 steps)
WM Forger 50 steps
WM Forger 100 steps
PixelSeal
0.025 / 6.4
0.003 / 0.8
0.533 / 136.3
0.280 / 71.8
Appendix
Table 6: Capacity of the watermarking systems against recent watermarking erasure attacks. The results are computed over 200 MFlickr 1024×1024 images.
Method
Rate Rσ / Rσ×M′
Identity
Brightness +0.2
Contrast ×2
JPEG QF =80
JPEG QF =50
PixelSeal
0.939 / 240.5
0.899 / 230.1
0.709 / 181.6
0.928 / 237.5
0.876 / 224.4
VideoSeal
0.878 / 224.7
0.793 / 203.1
0.539 / 138.1
0.863 / 221.0
0.816 / 209.0
TrustMark
0.611 / 61.1
0.530 / 53.0
0.309 / 30.9
0.545 / 54.5
0.491 / 49.1
SNW
0.983 / 686.0
0.868 / 666.3
0.514 / 394.6
0.884 / 679.0
0.837 / 642.5
Gaussian Blur 3×3 , σ=1
Rotation 90∘
Horizontal Flip
Hue 0.5
Saturation 1.5
Appendix
Table 7: Capacity of the watermarking systems against an extensive set of classic image transformations. The results are computed over 1000 MFlickr 1024×1024 images.
Figure 6: Example of an image watermarked with different methods and associated residuals for a fixed watermark power of 48 dB PSNR.
With the rapid proliferation of generative models, such as diffusion models, digital watermarking has emerged as a crucial solution for identifying AI-generated images. Modern post-hoc watermarking schemes use neural networks to achieve an extremely low false-alarm rate while remaining robust to common image transformations. However, there is a lack of comparison between these modern methods and classic ones, particularly in real-world scenarios where robustness and security take precedence over achieving an extremely low false-alarm probability. In this paper, we propose a fair comparison of robustness and security between modern and classic post-hoc watermarking across various types of classic augmentations and recent sophisticated attacks. Our experiments show that, in a realistic scenario, classic watermarking outperforms modern techniques in terms of security while maintaining robustness.
The rapid emergence of generative image models has led to the development of specialized watermarking techniques, particularly in-generation methods such as seed-based embedding. However, current evaluations in this area remain largely empirical, making them heavily reliant on the specific model architectures used for generation and inversion. This prevents any clear conclusion on the performance of any method, especially regarding security, for which a rigorous definition is lacking. Against this approach, we argue that the effectiveness of a watermarking scheme should be established purely through a thorough theoretical analysis. This is enabled by decoupling the model-dependent part from the actual decision mechanism of the watermarking system. Using this decoupling, we introduce a formal evaluation framework based on security, robustness, and fidelity. This allows precise comparisons between watermarking systems through a characteristic surface representing the trade-off between these three quantities, independent of any generative model. Based on this framework, we propose SSB, a novel watermarking method that generalizes previous seed-based methods by allowing to reach any security-robustness-fidelity regime on its characteristic surface. This work opens the door to the design of modern watermarking systems with theoretical guarantees that do not necessitate any costly empirical evaluations.
Invisible image watermarks are increasingly used for deepfake detection and provenance tracking, where they must survive not only incidental distortions but also deliberate removal. We revisit spread-spectrum embedding, a classical watermarking principle, inside a modern neural post-hoc watermarking architecture. Our starting point is a measurement: in existing encoder-decoder schemes each message bit occupies only a small fraction of the image, a shared contributing factor to their fragility, since removal then need only disturb the region a bit occupies. SpreadMark instead spreads each bit as a dense pseudo-random codeword over the whole image and recovers it by matched-filtering a learned cover-suppressed chip representation, with a parallel convolutional decoding path and sparsification-aware training. A conditional chip-space analysis shows that, under a codeword-independent perturbation model, dense spreading increases the budget required to disrupt matched-filter recovery. Evaluated on COCO and DIV2K against nine schemes, SpreadMark is the only evaluated method retaining high detection under both the regeneration and the latent-space sparsification settings we test, with competitive JPEG and additive-noise robustness. It keeps the embedded watermark imperceptible, maintaining high perceptual quality on both COCO and DIV2K.
Wei Song, Yuxin Cao, Zhenchang Xing +4
University of New South Wales, Australia · National University of Singapore, Singapore · CSIRO’s Data61, Australia