Ordinary prompt-based editing can cause latent watermark detection to fail without explicitly targeting the watermark. We benchmark eight watermark methods against five editors across four generative backbones, four editing strengths, and five semantic categories, with edit-validity and threshold checks. Separating editing from seven subsequent distortions reveals that editing alone primarily distinguishes Tree-Ring, while added distortions expose a broader spectrum of detection survival. Sequential edits reveal a second hidden difference: score separation can decline while detection rates remain near their ceiling. Across methods, standardized clean score separation (d′) organizes composite-survival tiers, whereas spatial overlap adds little to predicting edit-only survival beyond clean detectability. Embedding-strength interventions in two methods link higher clean separation to higher post-edit separation. In HSTR, the margin contrast is positive, while the angular layout contrast at matched clean separation remains unresolved. Together, outcome decomposition and continuous separation expose differences hidden by aggregate TPR. Method tiers are stable under threshold recalibration at the main operating points and alternative composite weights. Clean d′ is thus a useful empirical diagnostic within this benchmark, with mixed transfer to unseen methods. Code and supporting artifacts are planned for a separate release.
Figures & tables
Figure 1: Watermark survival under generative editing. Left: InfEdit Style Transfer examples; d0′ summarizes each method’s source bank. Right: a separate cumulative InfEdit cascade (Source, Object Swap, Style Transfer, Background Swap), showing TPR before further distortion, mean TPR over seven separately applied distortions, and score separation. The TPR plots include all eight methods; d′ is available for seven, with GS shown as N/A. Source counts and the five-step extension are reported in Appendix D .
Figure 2: Benchmark and analysis framework. Eight watermark methods are evaluated under five editors, four strength settings, and five edit categories. Detection is measured before and after seven separately applied distortions. Low/Medium/High/Max and the deviation arrow indicate the intended progression of each editor’s control settings; Max is the strongest setting tested.
P2P
P2P+NTI
InfEdit
FlowEdit-SD3
FlowEdit-FLUX
Method
Edit only
Post-dist.
Edit only
Post-dist.
Edit only
Post-dist.
Edit only
Post-dist.
Edit only
Post-dist.
Tree-Ring
0.632
0.356
0.804
0.471
0.682
0.444
0.692
0.403
0.734
0.427
ZoDiac
0.996
0.818
1.000
0.873
0.992
0.908
0.986
0.877
0.992
0.880
PRC
0.994
0.800
1.000
0.861
1.000
0.928
1.000
0.861
1.000
0.848
HSTR
1.000
0.934
1.000
0.952
1.000
0.965
1.000
0.944
1.000
0.937
HSQR
1.000
0.991
1.000
0.996
1.000
0.997
1.000
0.992
1.000
0.991
Table 1: Watermark survival at the main editing operating points. Entries are Avg 5 TPR over five edit categories. Edit only precedes added distortion; shaded Post-dist. columns average seven distortions applied separately to that output. P2P, P2P+NTI, InfEdit, and FlowEdit-SD3 use s=0.4 ; FlowEdit-FLUX uses s=0.2 , nmax=23 . Section E.1 lists native editor settings.
Figure 3: Edit-only TPR (top) and the mean TPR over seven separately applied post-edit distortions (bottom), across eight methods, five editors and four settings. Each point averages five categories equally. Higher s is closer to reconstruction within an editor. Table 9 retains the corresponding 8-condition composite sweep.
Editor
LPIPS
DINO
Δ CLIP
Δ BLIP2 ITM
P2P
0.282
0.793
0.020
0.177
P2P+NTI
0.283
0.747
0.025
0.202
InfEdit
0.250
0.751
0.034
0.308
FlowEdit-SD3
0.205
0.751
0.044
0.321
FlowEdit-FLUX
0.187
0.812
0.018
0.064
Table 2: Edit magnitude and target alignment relative to the watermarked input at the main operating points. Top: each editor aggregates Tree-Ring, ZoDiac, HSTR, HSQR, and GS over five categories and 100 source captions. LPIPS is a distance, DINO a similarity, and the two deltas measure changes in target alignment; Section E.2 defines the metrics. Bottom: a selected HSQR Object Swap example (source 59), shown for illustration rather than success-rate estimation.
Figure 4: Spatial diagnostics. (a) Mean source sensitivity ( n=100 per method); cyan guides mark the HSTR/HSQR latent-crop extent. Normalized maps do not compare absolute sensitivity. (b) Random-to-high-saliency mean score-drop ratios; 1 means equal mean damage. GS/TAG use STE maps and a separate log axis. Low-saliency results remain in Figure 8 . (c) The uniform IoU reference is geometric; Δ log-loss adds overlap to clean d′ ( − is better).
Figure 5: (a) Clean d′ versus 8-condition composite TPR with the global probit summary; circles denote fitted methods and diamonds held-out methods. (b) Strength intervention: clean versus mean category-level pre-distortion post-edit d′ , with marginal 95% bootstrap intervals. The arrow marks increasing α along each series. (c) HSTR effects on separation of source-mean scores, with paired-complete 95% bootstrap intervals. Left rows give matched clean targets 5/8; right rows fix Compact (C) or Distributed (D). The two effect axes have different ranges. All six layout intervals include zero; all six margin intervals are positive. Sections B.4 , B.5 and H.2 give the estimators.
Pooled model
AUC ↑
LPIPS only
0.633
Clean d′ only
0.916
d′ + LPIPS
0.945
Table 3: Per-image edit-only logistic prediction: pooled AUC ↑ (left) and held-out log-loss ↓ (right). Macro averages Tree-Ring and ZoDiac; the other four held-out methods are near ceiling. Section E.5 specifies the models.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Unit
Extension 500
Full 600
Tree-Ring
image
0.017
0.016
ZoDiac
image
0.020
0.018
HSTR
image
0.027
0.024
HSQR
image
0.028
0.025
PRC
image
0.034
0.030
T2S
key–image
0.010
0.010
Appendix
Table 4: Realized FPR under the original calibration thresholds, averaged over eight conditions. Extension 500 excludes the 100 threshold-setting images; Full 600 includes both cohorts. T2S and TAG pool 100 keys per image, so their key–image decisions are not independent images.
WM
Cond
Pooled FPR
Min key
Max key
Overdisp.
p
T2S
Clean
0.0100
0.0033
0.0200
0.915
0.7153
T2S
Brightness
0.0100
0.0017
0.0233
1.024
0.4157
T2S
Contrast
0.0100
0.0033
0.0250
0.915
0.7153
T2S
JPEG
0.0100
0.0017
0.0200
0.996
0.4909
T2S
Blur
0.0100
0.0017
0.0233
1.061
0.3196
T2S
BM3D
0.0100
0.0000
0.0233
1.153
0.1418
Appendix
Table 5: Per-key FPR for T2S and TAG. Overdispersion >1 describes variability above the reference level; statistical evidence for key dependence is assessed by the chi-squared p value. † marks p<0.01 .
Method
P2P
P2P+NTI
InfEdit
FE-SD3
FE-FLUX
Tree-Ring
0.25
0.36
0.32
0.29
0.32
ZoDiac
0.75
0.82
0.86
0.83
0.83
HSTR
0.93
0.95
0.97
0.95
0.94
HSQR
0.99
0.99
1.00
0.99
0.99
GS
1.00
1.00
1.00
1.00
0.99
Appendix
Table 6: Sensitivity to stricter operating rules (8-condition composite Avg 5 ): empirical zero-FP for four score-based methods and a stricter analytic rule for GS, as specified in the text.
Editor
Method
Avg 5 ( n=100 )
Avg 5 ( n=600 )
Δ
P2P
Tree-Ring
0.39
0.37
-0.024
ZoDiac
0.84
0.83
-0.012
PRC
0.82
0.80
-0.022
HSTR
0.94
0.90
-0.041
HSQR
0.99
0.99
+0.001
TAG
1.00
1.00
+0.000
Appendix
Table 7: Seven-method recalibration sensitivity in 8-condition composite Avg 5 TPR at the five main operating points: original q99 thresholds from 100 calibration images versus thresholds recalibrated on 600 images. Δ is the unrounded 600-bank minus 100-bank rate.
Method
P2P
P2P+NTI
InfEdit
FE-SD3
FE-FLUX
Clean d′
Tree-Ring
0.39 / 0.37
0.51 / 0.49
0.47 / 0.40
0.44 / 0.39
0.47 / 0.46
4.05
ZoDiac
0.84 / 0.82
0.89 / 0.87
0.92 / 0.90
0.89 / 0.88
0.89 / 0.89
8.03
PRC
0.82 / 0.79
0.88 / 0.84
0.94 / 0.93
0.88 / 0.87
0.87 / 0.86
9.62
HSTR
0.94 / 0.93
0.96 / 0.95
0.97 / 0.96
0.95 / 0.94
0.94 / 0.94
10.69
HSQR
0.99 / 0.99
1.00 / 0.99
1.00 / 0.99
0.99 / 0.99
0.99 / 0.99
20.67
TAG
1.00 / 1.00
1.00 / 1.00
1.00 / 1.00
1.00 / 1.00
1.00 / 0.99
22.60 †
Appendix
Table 8: 8-condition composite at the main editing operating point. Each cell reports Avg 5 / worst over five edit categories after first averaging the Edit-only result and seven separately applied subsequent distortions within each category. This composite is retained for comparison with the full strength sweep and clean- d′ analyses; it is distinct from Edit only and Post-distortion mean in Table 1 . † TAG’s clean d′ is finite but uninformative under clean-score saturation; GS’s “—” denotes d′ not computed.
P2P
P2P+NTI
InfEdit
Method
0.7
0.4
0.2
0.0
0.7
0.4
0.2
0.0
0.7
0.4
0.2
0.0
Tree-Ring
0.396
0.391
0.381
0.384
0.516
0.513
0.504
0.501
0.562
0.474
0.446
0.442
ZoDiac
0.843
0.841
0.831
0.823
0.893
0.888
0.876
0.881
0.930
0.918
0.912
0.903
PRC
0.832
0.824
0.815
0.806
0.889
0.878
0.878
0.869
0.934
0.937
0.931
0.932
HSTR
0.939
0.942
0.930
0.925
0.962
0.958
0.955
0.958
0.975
0.970
0.970
0.971
HSQR
0.992
0.992
0.991
0.993
0.997
0.997
0.996
0.994
0.997
0.997
0.996
0.996
Appendix
Table 9: Full sweep of the 8-condition composite Avg 5 TPR: eight watermarks, five editors, and four settings (160 cells). Each cell averages the Edit-only TPR and seven separately applied post-edit distortion TPRs within category before averaging five categories. Higher s is closer to reconstruction within an editor, but s is not a common cross-editor severity scale. Bold marks the main-table setting (FlowEdit-FLUX: s=0.2 ; all others: s=0.4 ).
Budget ( nmax )
LPIPS
DINO
Δ CLIP
Δ BLIP2 ITM
nmax=17 (nominal)
0.089
0.928
0.0031
-0.024
nmax=23 (selected)
0.187
0.812
0.0177
0.064
nmax=28
0.538
0.522
0.0425
0.275
SD3 (reference)
0.205
0.751
0.0438
0.321
Appendix
Table 10: FlowEdit-FLUX edit validity across solver budgets, relative to the watermarked editing input. Each row aggregates five original methods, five categories, and 100 source captions (2,500 pairs). LPIPS is a distance and DINO a similarity, so their magnitudes indicate change rather than intrinsic improvement. Δ CLIP and Δ BLIP2 ITM report changes in target alignment.
Held-out
d′
Position
Mean abs. err
Max abs. err
Tree-Ring
4.05
extrap.(low)
0.2622
0.3329
ZoDiac
8.03
interp.
0.0785
0.1923
PRC
9.62
interp.
0.0692
0.1395
HSTR
10.69
interp.
0.0120
0.0472
HSQR
20.67
interp.
0.0087
0.0419
T2S
23.35
extrap.(high)
0.0031
0.0204
Appendix
Table 11: Leave-one-method-out evaluation of condition-specific probit fits for 8-condition composite TPR. Mean and maximum absolute prediction errors summarize the 20 editor–strength conditions, in TPR units. These are separate from the global curve and the per-image severity logistic models.
Figure 6: HSTR embedding strength intervention. The central 44×44 crop of the initial 64×64 latent is transformed, and masked coefficients are blended as Fα[m]=(1−α)F[m]+αW[m] before inverse transformation and writeback. The masks and pattern magnitudes are shown separately for channels 0 and 3, using a shared magnitude scale without clipping. The RGB strip shows outputs for caption index 60; the unwatermarked reference is a separate generation with the same caption.
Method
α
LPIPS
PSNR
SSIM
HSTR
0.15
0.2827
18.19
0.6969
0.30
0.3847
15.89
0.5822
0.50
0.4499
14.55
0.5108
0.70
0.4834
13.67
0.4707
1.00
0.5170
12.47
0.4268
Tree-Ring
0.15
0.2545
19.30
0.7297
Appendix
Table 12: Source fidelity across embedding strengths. LPIPS uses VGG at 512×512 relative to the unwatermarked generation; SSIM is the source script’s simple implementation. Entries are descriptive means over n=100 images per row.
Figure 7: Illustrative masking examples for HSTR (top) and HSQR (bottom) on the source for caption index 0, with p=20% of the central 352×352 ROI (dashed box), i.e., k=24,781 pixels per mask. Each block shows the source, detection saliency with the ROI, and high-saliency, random (replication 0; the five-draw mean is noted), and low-saliency masks with the resulting detector inputs. White marks masked pixels, which are replaced by the source’s channel-wise mean RGB. Displayed drops are stored unmasked-minus-masked native detector scores for these examples, not changes in d′ or aggregate random-mask results. These illustrations do not establish targeted–random equivalence or locate the physical watermark signal; Figures 8 and 13 give the aggregate comparison.
Figure 8: Masking ablation on HSQR and HSTR under the support-matched protocol. Targeted (high-saliency) and random masking have similar damage at larger fractions. Low-saliency masking is not uniformly less damaging: HSTR has a larger low- than high-saliency drop at p=30% and 50% . Random masking is averaged over five draws per source; ∼ marks the stored TOST result with a margin of ±10% of the mean targeted drop.
Method
p %
ratio (rnd/high)
dz
TOST p
Equiv?
HSQR
5
1.166
-0.752
0.9983
no
10
1.036
-0.211
0.0001
yes
20
0.953
0.353
0.0001
yes
30
0.925
0.655
0.0139
yes
50
0.924
0.787
0.0071
yes
HSTR
5
1.227
-1.265
1.0000
no
Appendix
Table 13: Masking ablation: random-to-targeted score-damage ratio and the reported TOST under a 10% margin. The margin is not independently justified, so the “Equiv?” column is a protocol-specific sensitivity result rather than a general equivalence claim.
Method
p=5%
p=10%
p=20%
p=30%
p=50%
ZoDiac
1.22
1.12
1.00
0.97
0.95
PRC
1.23
1.10
1.02
1.00
1.00
T2S
1.43
1.24
1.07
1.00
0.97
GS
80.8
20.1
4.56
2.53
1.41
TAG
5.60
3.11
1.91
1.45
1.20
Appendix
Table 14: Extended masking ablation: random-to-targeted damage ratio for five additional watermarks. Ratios generally move toward one as the mask fraction increases, but GS and TAG remain above one even at p=50% . Their extreme small- p ratios are interpreted as secondary evidence because clean bit accuracy is saturated and saliency uses an STE surrogate (see text).
Method
Clean d′
Gini
Entropy n
Top-10%
Moran’s I
Center
Radial
frac
ratio
HF/LF
Tree-Ring
4.05
0.605
0.939
0.455
0.630
0.419
0.0026
ZoDiac
8.03
0.605
0.940
0.454
0.659
0.402
0.0021
PRC
9.62
0.656
0.924
0.507
0.634
0.409
0.0022
HSTR
10.69
0.662
0.919
0.514
0.622
0.901
0.0015
HSQR
20.67
0.629
0.927
0.484
0.635
0.882
0.0014
Appendix
Table 15: Clean separation and six spatial distribution statistics for 8 watermarks (support mask, n=100 images). † : TAG’s finite clean d′ is secondary under score saturation; GS has no compatible clean d′ estimate. The reported statistics show no monotone ordering aligned with available clean d′ or 8-condition composite TPR.
Figure 9: Detection saliency before and after an InfEdit Object Swap for Tree-Ring, ZoDiac, and HSQR, alongside cross-attention differences. Rows use method-specific source images with a shared caption index. Each d0′ describes a source bank. Saliency maps visualize detector sensitivity; attention maps summarize changes during editing.
Figure 10: HSQR under InfEdit across five semantic edit categories. Detection saliency is evaluated on the source and edited inputs; cross-attention difference summarizes the editing process.
Figure 11: Additional InfEdit examples for caption indices 0, 6, and 15 (idx). Rows show Tree-Ring, ZoDiac, HSTR, and HSQR, each with its own watermarked source for the same caption. The first two columns show the source and its detection saliency; each edit category (Style Transfer, Object Swap, Attribute/Color Change) then pairs the edited image with its cross-attention difference. Along a row, categories are compared within one method; down a column, methods are compared under the same edit.
Figure 12: Tree-Ring and HSQR under FlowEdit-SD3 ( s=0.2 ). The final column shows RGB pixel changes between the source and edited image, alongside the verified detection outcome for each displayed case.
Figure 13: HSQR under InfEdit across five editing categories and four editing strengths ( s=0.7,0.4,0.2,0.0 ). All twenty edited images maintain successful detection, consistent with HSQR’s near-saturated robustness in the main benchmark.
Step 0
Step 1
Step 2
Step 3
Step 4
Step 5
Method
TPR
d′
TPR
d′
TPR
d′
TPR
d′
TPR
d′
TPR
d′
Tree-Ring
0.98
4.05
0.77
2.83
0.40
1.86
0.21
1.48
0.12
1.18
0.06
0.92
ZoDiac
1.00
8.03
0.99
5.31
0.95
3.87
0.93
3.19
0.87
2.84
0.77
2.44
PRC
1.00
9.62
1.00
5.54
1.00
4.38
1.00
3.91
1.00
3.40
1.00
3.09
HSTR
1.00
10.69
1.00
6.48
1.00
5.32
1.00
4.97
1.00
4.43
0.98
3.92
HSQR
1.00
20.67
1.00
11.65
1.00
9.86
1.00
9.17
1.00
8.45
1.00
7.82
Appendix
Table 16: Cumulative InfEdit trajectory: Step 0 is the source; Steps 1–5 add Object Swap, Style Transfer, Background Swap, Attribute/Color Change, and Pose/Shape Change. Each edit takes the preceding output. Cross/self-attention is 0.4 ; denoising strength is 1.0 with 15 inference steps. Values precede further distortion. Source n=100 ; Tree-Ring and HSQR use 99 at Steps 1–4, then 98 and 99 at Step 5; all other rows use 100. GS d′ is unavailable.
Figure 14: Watermarked and null score distributions along the cumulative InfEdit sequence of Table 16 . Rows show the seven methods with an empirical null; columns show the source (Step 0) and Steps 1–5 (Object Swap, Style Transfer, Background Swap, Attribute/Color Change, Pose/Shape Change) before further distortion. Gray: the unedited, non-watermarked source-null bank that sets each threshold ( Section A.1 ); color: watermarked scores. Scores are in native units, oriented so that larger values indicate stronger evidence, with axes shared within each row. Scores outside the plotted range are omitted: at most six watermarked scores per row, all at Step 0, and under 1% of T2S and TAG null pairs. Dashed lines mark the q99 threshold; labels give d′ and TPR from Table 16 . Triangles mark bars truncated at the row’s y -limit (TAG’s source and Step-1 watermarked scores concentrate near bit accuracy 1.0 ). GS is omitted because it has no compatible empirical null.
Method
Step
n
No dist.
Bright.
Contr.
JPEG
Blur
BM3D
VAE-B
VAE-C
Mean 7
Comp. 8
Tree-Ring
0
100
0.980
0.460
0.930
0.500
0.920
0.820
0.510
0.510
0.664
0.704
1
99
0.768
0.323
0.687
0.404
0.737
0.576
0.424
0.374
0.504
0.537
2
99
0.404
0.141
0.343
0.283
0.364
0.273
0.333
0.293
0.290
0.304
3
99
0.212
0.091
0.152
0.182
0.202
0.172
0.242
0.192
0.176
0.181
4
99
0.121
0.040
0.081
0.141
0.111
0.091
0.172
0.131
0.110
0.111
5
98
0.061
0.020
0.051
0.092
0.071
0.061
0.143
0.102
0.077
0.075
Appendix
Table 17: Detection after cumulative InfEdit editing, for all eight methods and every subsequent distortion. Each distortion is applied separately to the corresponding cumulative output. Step 0 is the source; Steps 1–5 add Object Swap, Style Transfer, Background Swap, Attribute/Color Change, and Pose/Shape Change. No dist. is the output before further distortion; Mean 7 excludes that column, whereas Comp. 8 includes it. Bright. and Contr. denote brightness and contrast. n is the number of source identities in each row. Empirical q99 scores use strict > ; GS uses native bit accuracy ≥147/256 .
Figure 15: Examples of the first three edits in the cumulative InfEdit procedure: Source, Object Swap, Style Transfer, and Background Swap. Every edit takes the preceding output as input. The eight method rows share caption index 60 and use method-specific source images. The displayed prompt changes describe the requested edits; the images show the resulting transformations, including retained background content. The additional fourth and fifth edits are reported numerically in Tables 16 and 17 .
Method
Category
Mean TPR
95% CI
Tree-Ring
Attribute_Color Change
0.4044
± 0.0077
Background Swap
0.3931
± 0.0444
Object Swap
0.3978
± 0.0216
Pose_Shape Change
0.4153
± 0.0071
Style Transfer
0.4137
± 0.0115
ZoDiac
Attribute_Color Change
0.9000
± 0.0097
Appendix
Table 18: Verification-TPR intervals from a separate repeated-seed cohort for all eight methods. For each method–category pair, TPR is first averaged over the eight post-edit conditions within each of five disjoint seed namespaces (1042, 2042, 3042, 4042, 5042); the table gives the five-namespace mean and 95% Student- t half-width (four degrees of freedom). Retained counts are 95–100 for the original five methods and 96–100 for PRC, T2S, and TAG per namespace and category, so these are not intervals for every main benchmark cell.
Editor
Backbone
Procedure
Guidance
Parameter
0.7
0.4
0.2
0.0
P2P
SD2.1-base
DDIM inversion (empty prompt, 50 steps), then 50-step DDIM editing
7.5
τcross=τself
0.7
0.4
0.2
0.0
P2P+NTI
SD2.1-base
Null-text inversion (50 steps, 10 inner steps), computed once per source; editing as in P2P
7.5
τcross=τself
0.7
0.4
0.2
0.0
InfEdit
LCM DreamShaper v7
Inversion-free, 15 LCM steps
2.0 (source 1.0)
τcross=τself
0.7
0.4
0.2
0.0
FlowEdit-SD3
SD3 Medium
T=50 , nmin=0
3.5 / 13.5
nmax
15
30
40
50
FlowEdit-FLUX
FLUX.1-dev
T=28 , nmin=0
1.5 / 5.5
nmax
9
17
23
28
Appendix
Table 19: Native editor settings at s∈{0.7,0.4,0.2,0.0} ; bold marks the main operating point. FlowEdit guidance is source/target; InfEdit’s Pose/Shape edits use τcross=0.6 . All editors seed each source with 42+idx , shared across categories. P2P and P2P+NTI apply attention replacement or refinement according to the edit type, with local blending on the edited words except in style and pose edits; InfEdit applies attention refinement, with local blending on the target keyword in background, object, and attribute edits.
Method
Editors
Unadjusted range
Adjusted range
∣Δ∣max
Tree-Ring
5
0.594–0.839
0.582–0.838
0.014
ZoDiac
5
0.986–1.000
0.985–1.000
0.001
PRC
3
0.995–1.000
0.993–1.000
0.002
HSTR
5
1.000–1.000
1.000–1.000
0.000
HSQR
5
1.000–1.000
1.000–1.000
0.000
T2S
3
1.000–1.000
1.000–1.000
0.000
Appendix
Table 20: Edit-only TPR on common LPIPS support, pooling all four strengths and five categories. Ranges span P2P, P2P+NTI, InfEdit, FlowEdit-SD3, and FlowEdit-FLUX for Tree-Ring/ZoDiac/HSTR/HSQR, and the first three editors for PRC/T2S. ∣Δ∣max is the largest adjusted-minus-unadjusted magnitude on the same retained cohort. Tree-Ring’s 0.594 – 0.839 range differs from Table 1 ’s 0.632 – 0.804 because this cohort pools strengths and restricts LPIPS, whereas that table uses the main points.
Method
Embedding
Key / payload
Detection score
Tree-Ring
Ring pattern in the channel-3 Fourier disk of zT (radius 14; 613 coefficients)
Per-image key from a 2,048-key codebook
−L1 distance
HSTR
Central 44×44 latent crop: Hermitian-symmetric ring pattern in the channel-3 disk (613 coefficients) and a Gaussian key in the channel-0 ring (556)
Per-image key from a 2,048-key codebook
− minimum channel L1
HSQR
QR code (version 1) in channel-3 rFFT coefficients of the central crop (1,764)
Per-image key from a 2,048-key codebook
−L1 distance
ZoDiac
Ring pattern in the channel-3 Fourier disk (radius 10; 317 coefficients) of an optimized latent (100 Adam iterations, SSIM target 0.92)
One key
−L1 distance
GS
256-bit message over the 4×64×64 latent ( 8×8 spatial replication), ChaCha20-encrypted
Per-image key and message
Bit accuracy
PRC
Pseudorandom code over 16,384 latent signs
One global key; 512-bit message
Parity log-likelihood
Appendix
Table 21: Watermark settings. Each method uses the default embedding settings of its reference implementation; Tree-Ring, HSTR, and HSQR share one Fourier-watermark implementation. Null images are scored with the same keys as watermarked images.
Figure 16: HSTR angular layout intervention in the Fourier coordinates of the central 44×44 latent crop. Compact and Distributed each select 307 channel 3 coefficients with identical radial pair counts; the 556-coefficient channel 0 ring is fixed. Separate α values achieve clean d′ near 5 or 8 on 100 sources per condition. Editing covers 100 sources × 5 categories × 3 editors × 4 conditions: 6,000 nominal and 5,973 available outputs.
Editor
Comparison
Target d′
n
Effect
95% CI
0 in CI?
P2P+NTI
Layout
5
100
−0.008
[−0.700,+0.680]
yes
P2P+NTI
Layout
8
100
−0.144
[−0.996,+0.679]
yes
InfEdit
Layout
5
97
+0.083
[−0.620,+0.818]
yes
InfEdit
Layout
8
94
+0.097
[−0.820,+1.050]
yes
FlowEdit-SD3
Layout
5
100
+0.096
[−0.506,+0.701]
yes
FlowEdit-SD3
Layout
8
100
+0.151
[−0.499,+0.806]
yes
Appendix
Table 22: Angular layout and margin effects from paired complete sources ( Section H.2 ). Layout is Compact minus Distributed at matched target d′ ; margin is target d′≈8 minus d′≈5 at fixed layout. n counts sources with all five categories available in both conditions. The 5,000-iteration percentile bootstrap jointly resamples edited-source IDs and independently resamples the two null arrays. “0 in CI” indicates an unresolved effect.
This paper investigates a fundamental yet underexplored question: can watermarked images remain editable without compromising watermark integrity? We propose SafeMark, a framework for watermark-preserving text-guided image manipulation that explicitly integrates watermark integrity into the editing process. Specifically, SafeMark adds a thresholded watermark-decoding loss directly to the diffusion editor's training objective, fine-tuning the editor so that semantically valid edits also preserve the embedded watermark at the final output. This design admits a clean information-theoretic justification: maintaining high bit-accuracy on the edited image lower-bounds the mutual information that the editor channel preserves between watermark and edited output, the quantity that fundamentally controls watermark recoverability. SafeMark is compatible with differentiable diffusion-based editors, and requires no architectural modification. Extensive evaluations across multiple datasets, text-guided editing methods, and post-edit distortion settings demonstrate that SafeMark achieves high watermark bit accuracy across diverse editing settings while maintaining high-quality semantic edits, without sacrificing robustness to common post-edit distortions. These results demonstrate that semantic editability and watermark integrity are fundamentally compatible, enabling trustworthy image provenance in generative editing pipelines.
Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adversary to alter critical content while retaining attribution, a vulnerability known as piggyback spoofing. We introduce an innovative watermark that jointly provides provenance and tamper evidence. It co-embeds a robust signal and a fragile signal into each generated token. The signals share the same mechanism but use independent keys and different seeding windows over normalized text, making one resilient to edits and the other sensitive to reader-visible changes. Multiple rounds of unbiased tournament reweighting preserve the expected generation distribution, while a periodic round-allocation pattern controls the trade-off between the two signals. At detection, their scores form a two-dimensional space supporting three decisions: Intact, Tampered, and No-Watermark. Across two large language models and two prompt datasets, our method demonstrates the highest tamper-detection rate among the evaluated methods while maintaining competitive attribution robustness and perplexity. Ablation studies show that reliable three-state detection requires a well-defined notion of intactness, co-embedding of the two signals, and complementary sensitivity to edits.
With the growing use of large language models, concerns over content authenticity have spurred a variety of watermarking schemes. These schemes use secret keys to detect machine-generated text while remaining imperceptible to readers. Detection typically reduces to statistical hypothesis testing for the presence of watermarks, a topic that is now well studied. In contrast, the finer-grained task of localizing which segments of a text are watermarked is much less explored; existing approaches often lack scalability or guarantees robust to paraphrasing and post-editing. We bring a new perspective to this segmentation problem through the lens of epidemic change-points and, by exploiting this connection, propose WISER, a novel and computationally efficient watermark segmentation algorithm. We establish finite-sample error bounds and consistency for detecting multiple watermarked segments in a single text. Complementing these theoretical results, our extensive numerical experiments show that WISER outperforms state-of-the-art baseline methods, both in terms of computational speed as well as accuracy, on various benchmark datasets embedded with diverse watermarking schemes. Together, these theoretical and empirical results position WISER as an effective tool for watermark localization and illustrate how classical statistical ideas can yield theoretically valid and computationally efficient solutions to a modern problem of immediate importance.
Soham Bonnerjee, Subhrajyoty Roy, Sayar Karmakar
University of Chicago · Washington University in St. Louis · University of Florida