The proliferation of high-fidelity generative editing models has made it possible to inject violent or sexual content into otherwise ordinary images while preserving visual plausibility, with concrete consequences for public discourse and vulnerable populations. We propose a robust semantic watermarking framework that reframes the watermark as a recoverable semantic reference rather than an opaque identifier. Our framework combines a β-VAE-based binary watermark (CLIP-VAE) with explicit channel-aware training---random bit-flip noise is injected during training so that the decoder learns graceful degradation under the noisy watermarking channel. As a downstream application, a lightweight module SDA-Net uses the recovered semantic embedding to expose not only whether but in which semantic direction an image has been altered. In a 5-way comparison against representative binary hashing baselines (SimHash, ITQ, HashNet, and their robust-MLP variants), CLIP-VAE achieves the highest reconstruction cosine similarity to the original CLIP embedding under realistic InstructPix2Pix bit-error rates, and uniquely supports direction-of-drift detection---a forensic complement to existing content-moderation pipelines.
Figures & tables
Category
Mean
Std.
Range
Normal
0.7948
0.0601
[0.5120, 0.9496]
Violence
0.8409
0.0642
[0.5678, 0.9552]
Sexual
0.8832
0.0446
[0.7165, 0.9540]
Overall
0.8285
0.0685
[0.6943, 0.9626]
Table 1: Semantic preservation performance on the test set measured by cosine similarity between original and reconstructed CLIP embeddings.
Class
Precision
Recall
F1 Score
Normal
0.9765
0.9875
0.9820
Violence
0.9747
0.9626
0.9686
Sexual
0.9949
0.9850
0.9899
Overall
0.9821
0.9784
0.9802
Table 2: Per-class classification performance of SDA-Net on the test set.
k flipped
5
10
20
30
SimHash + MLP (robust)
0.824
0.802
0.734
0.634
ITQ + MLP (robust)
0.830
0.812
0.763
0.698
HashNet
0.644
0.625
0.586
0.545
CLIP-VAE (base)
0.807
0.791
0.747
0.683
CLIP-VAE (ours)
0.830
0.819
0.783
0.717
Table 3: Bit-flip reconstruction cosine similarity in the realistic InstructPix2Pix BER regime (mean over 10 random-flip trials, std≤0.0024 ). Channel-aware CLIP-VAE achieves the highest score across all practical bit-error rates. k=0 (clean) and k=50 (near-random) are reported in Appendix D .
State
Pred.
Δlatent
ΔVio
ΔSex
Original
Normal
—
—
—
Watermarked
Normal
1.98
− 0.1
+ 0.2
Manipulated
Normal
4.65
− 2.3
+ 0.4
Table 4: Sequential drift analysis. Δlatent denotes drift magnitude from the original; Δc denotes per-class distance change (negative = closer to class c ). The predicted class remains Normal in all three states, but ΔVio=−2.3 reveals directional drift toward the Violence prototype before any classifier boundary is crossed.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Latent Dim
Latent Cosine ↑
CLIP Cosine ↑
Category Acc. ↑
50D
0.8070
0.9525
0.8877
100D
0.8055
0.9648
0.8942
Appendix
Table 5: Effect of latent dimensionality on semantic reconstruction.
β
CLIP Cosine Sim ↑
KL Divergence ↓
0.0
0.8392
5.47
0.001
0.8568
2.06
0.01
0.8339
0.70
0.1
0.7965
0.19
Appendix
Table 6: Effect of KL weight β on reconstruction quality and latent regularity.
Figure 4 : Latent space visualizations under different KL weights β . With β=0 , the latent is unconstrained and asymmetric; β=0.001 leaves the distribution insufficiently regularized for binarization; β=0.01 produces the desired well-distributed, class-separated latent geometry; β=0.1 over-regularizes and collapses class structure.
Model
Cosine Similarity ↑
Cosine Distance ↓
VAE
0.8366
0.1634
VQ-VAE
0.7677
0.2323
Appendix
Table 7: Comparison between VAE and VQ-VAE on semantic reconstruction quality.
k flipped
0 (clean)
50 (near-random)
SimHash + MLP (robust)
0.840
0.444
ITQ + MLP (robust)
0.845
0.537
HashNet
0.662
0.465
CLIP-VAE (base)
0.819
0.538
CLIP-VAE (ours)
0.838
0.536
Appendix
Table 8: Bit-flip reconstruction cosine at extreme k (clean and near-random).
Component
Setting
Purpose
Classification
λcls=1.0
CE + label smoothing 0.1
KL Divergence
λKL=0.01
Regularize to N(0,I)
SCL
λSCL=0.5
Class separation ( τ=0.07 )
Prototype EMA
m=0.9
Stabilize prototypes
Appendix
Table 9: SDA-Net training hyperparameters.
Variant
cos@k=30
LinProbe
neither
0.684
0.854
sign-margin only
0.689
0.900
flip-noise only
0.708
0.941
flip-noise + sign-margin (ours)
0.702
0.918
Appendix
Table 10: Component ablation of channel-aware training. flip-noise alone provides the bulk of the robustness benefit; sign-margin acts as a smaller complementary regularizer.
Configuration
cos@k=30
LinProbe
λflip=0.05
0.703
0.880
λflip=0.10
0.703
0.828
λflip=0.20
0.704
0.896
λflip=0.50
0.697
0.813
margin=0.1,kmax=5
0.704
0.872
margin=0.1,kmax=15
0.713
0.917
Appendix
Table 11: Hyperparameter sensitivity of channel-aware training. The robustness metric ( cos@k=30 ) is stable across all configurations; linear-probe accuracy shows mild variation.
This paper investigates a fundamental yet underexplored question: can watermarked images remain editable without compromising watermark integrity? We propose SafeMark, a framework for watermark-preserving text-guided image manipulation that explicitly integrates watermark integrity into the editing process. Specifically, SafeMark adds a thresholded watermark-decoding loss directly to the diffusion editor's training objective, fine-tuning the editor so that semantically valid edits also preserve the embedded watermark at the final output. This design admits a clean information-theoretic justification: maintaining high bit-accuracy on the edited image lower-bounds the mutual information that the editor channel preserves between watermark and edited output, the quantity that fundamentally controls watermark recoverability. SafeMark is compatible with differentiable diffusion-based editors, and requires no architectural modification. Extensive evaluations across multiple datasets, text-guided editing methods, and post-edit distortion settings demonstrate that SafeMark achieves high watermark bit accuracy across diverse editing settings while maintaining high-quality semantic edits, without sacrificing robustness to common post-edit distortions. These results demonstrate that semantic editability and watermark integrity are fundamentally compatible, enabling trustworthy image provenance in generative editing pipelines.
Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the watermark detection pipeline on neural networks introduces a critical yet underexplored backdoor attack surface. To systematically study this vulnerability, we propose GhostVAE to plant a stealthy backdoor into the encoder of Variational Autoencoder (VAE), enabling reliable evasion of watermark detection. GhostVAE operates in two stages: it first constructs a universal trigger via power spectrum regularization to improve the trigger robustness, and then trains a backdoored VAE encoder with a parameter-aligned objective. Through extensive evaluations across three state-of-the-art semantic watermarking schemes and three widely adopted LDMs, we show that GhostVAE preserves watermark detection performance on benign images (achieving an average true positive rate of 94.4%), while simultaneously enabling highly effective evasion under trigger activation (achieving an average attack success rate of 94.6%). Moreover, we comprehensively analyze seventeen representative defenses and demonstrate that GhostVAE remains stealthy across the input space, parameter space, and latent space. Our work fundamentally undermines the trustworthiness of semantic watermarking systems and highlights that secure deployment of semantic watermarks requires end-to-end security considerations, particularly for neural network components.
Jinyuan Liu, Tianshuo Cong, Pei Li +4
1Tsinghua University · 2Shandong University · 4Shandong Key Laboratory of Artificial Intelligence Security, Shandong University +5
The rapid progress of image-guided video generation (I2V) has raised concerns about its potential misuse in misinformation and fraud, underscoring the urgent need for effective digital watermarking. While existing watermarking methods demonstrate robustness within a single modality, they fail to trace source images in I2V settings. To address this gap, we introduce the concept of Robust Diffusion Distance, which measures the temporal persistence of watermark signals in generated videos. Building on this, we propose I2VWM, a cross-modal watermarking framework designed to enhance watermark robustness across time. I2VWM leverages a video-simulation noise layer during training and employs an optical-flow-based alignment module during inference. Experiments on both open-source and commercial I2V models demonstrate that I2VWM significantly improves robustness while maintaining imperceptibility, establishing a new paradigm for cross-modal watermarking in the era of generative video. \href{https://github.com/MrCrims/I2VWM-Robust-Watermarking-for-Image-to-Video-Generation}{Code Released.}