Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated sentences without explicitly accounting for the highly non-uniform and context-dependent semantic preference of LLM generation. When these two preferences are poorly aligned, many natural continuations become incompatible with the watermark, causing semantic narrowing: reduced semantic freedom, increased resampling cost, and potential degradation on tasks with strict semantic requirements. To alleviate this problem, we propose HammingMark, which uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood. Defining watermark validity over a Hamming neighborhood in compact hash space retains a larger fraction of naturally likely semantic continuations. The coarse many-to-one hash mapping further allows diverse semantic realizations to remain watermark-valid. Experiments on C4 and BookSum show that HammingMark achieves strong robustness, high detectability, and near-unwatermarked generation quality, requiring only 2.2 sampled candidates per accepted sentence,a 72.8% reduction compared with the most sampling-efficient existing method. On more complex tasks with strict semantic constraints, HammingMark achieves the highest detection rates with the highest or tied-highest ROUGE-L scores, demonstrating its effectiveness in balancing watermark detectability and generation quality under constrained generation settings.
Figures & tables
Method
Clean
Parrot
Pegasus
DIPPER
API
PPL ↓
Mistral-7B ( Jiang et al., 2023 ) on C4 ( Raffel et al., 2020 )
No watermark
–
–
–
–
–
4.51
KGW ( Kirchenbauer et al., 2023 )
1.00/1.00/1.00
0.59/0.72/0.90
0.51/0.63/0.89
0.43/0.49/0.84
0.31/0.39/0.80
8.00
SynthID ( Dathathri et al., 2024 )
1.00/1.00/1.00
0.32/0.45/0.82
0.23/0.42/0.80
0.35/0.40/0.80
0.24/0.38/0.79
4.52
MorphMark ( Wang et al., 2025 )
0.99/1.00/1.00
0.54/0.60/0.88
0.47/0.55/0.84
0.51/0.55/0.86
0.53/0.64/0.89
6.62
SIR ( Liu et al., 2024 )
0.98/1.00/1.00
0.51/0.63/0.91
0.56/0.60/0.90
0.44/0.57/0.84
0.38/0.46/0.82
8.35
Table 1: Clean detectability, robustness, and generation quality on C4 and BookSum. Each detection entry reports TPR@1%/TPR@5%/AUROC .
Method
SimMark
SemStamp
K-SemStamp
PMark
Random Code
HammingMark
Samples / sent. ↓
8.1
99.3
13.3
64.0
14.7
2.2
Tokens / sent. ↓
186.7
1694.4
246.9
1185.8
273.1
41.58
Natural acceptance αˉ↑
0.34
0.04
0.13
/
0.15
0.58
Table 2: Sampling efficiency and estimated natural acceptance on C4. Lower sampling cost and higher natural acceptance are better. Random Code uses the same code-space coverage as HammingMark ( 37/256 ) but selects valid codes uniformly at random.
Method
ELI5 ( Fan et al., 2019 )
Multi-News ( Fabbri et al., 2019 )
@1 ↑
@5 ↑
AUC ↑
Rouge-L ↑
@1 ↑
@5 ↑
AUC ↑
Rouge-L ↑
No watermark
–
–
–
0.34
–
–
–
0.36
KGW
0.78
0.89
0.97
0.29
0.21
0.49
0.86
0.31
SynthID
0.43
0.67
0.92
0.31
0.08
0.17
0.64
0.36
MorphMark
0.49
0.64
0.90
0.29
0.26
0.41
0.88
0.35
SIR
0.63
0.75
0.93
0.28
0.21
0.36
0.85
0.34
Table 3: Performance on complex generation tasks on Mistral-7B. Detection results are reported as TPR@1%, TPR@5%, and AUROC.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Reference assumption
Accepted
Excluded
SemStamp
Equal partition masses
25%
75%
K-SemStamp
Equal partition masses
25%
75%
Cosine-SimMark
Uniform angle; [a,b]=[0.68,0.76]
3.68%
96.32%
Random Code
Uniform selection of 37 codes; set average
14.45%
85.55%
HammingMark
Uniform angle; projection average
33.33%
66.67%
Appendix
Table 4: Idealized reference calculations, with margins omitted. The rows use different assumptions and averaging operations; they explain mechanisms rather than rank performance on one real generation distribution.
Method
Setting
Candidates/sentence
TPR@1% FPR
SemStamp
γ=0.5
2.9
6%
K-SemStamp
γ=0.5
3.1
9%
SimMark
[a,b]=[0.40,0.90]
1.6
7%
HammingMark
m=8,T=6
2.2
100%
Appendix
Table 5: Relaxed-constraint comparison.
Method
C1
C2
C3
C4
C5
C6
C7
C8
Total
SemStamp
1
0
0
0
0
0
0
0
1
K-SemStamp
1
1
0
0
0
0
0
0
2
SimMark
7
5
4
0
0
0
0
0
16
PMark
10
8
6
4
3
0
0
0
31
HammingMark
24
20
17
15
12
7
4
0
99
Appendix
Table 6: Distribution of watermark-valid candidates across eight semantic clusters.
Method
Pairwise diversity ↑
Diversity retention ↑
No watermark
0.284
1.000
SimMark
0.224
0.789
SemStamp
0.142
0.500
K-SemStamp
0.158
0.556
PMark
0.201
0.708
Random Code
0.184
0.648
Appendix
Table 7: Semantic diversity under fixed contexts. Each method independently generates 300 continuations for each of 100 prompts. Higher values indicate greater semantic diversity.
Dataset
Generator
Fallback rate (%)
C4
Mistral-7B
1.3
BookSum
Mistral-7B
0.8
C4
Qwen2.5-7B
1.2
BookSum
Qwen2.5-7B
0.6
ELI5
Mistral-7B
2.3
Multi-News
Mistral-7B
2.7
Appendix
Table 8: Empirical fallback rates under the maximum resampling budget B=16 .
M
T
Clean
Parrot
DIPPER
4
3
0.76/0.86/0.97
0.75/0.86/0.97
0.74/0.83/0.96
8
6
0.98/1.00/1.00
0.94/0.99/1.00
0.93/0.98/0.99
16
12
0.98/1.00/1.00
0.78/0.87/0.98
0.73/0.81/0.96
32
24
0.99/1.00/1.00
0.63/0.73/0.90
0.64/0.74/0.91
Appendix
Table 9: Effect of semantic hash length on clean detection and attack robustness. Each entry reports TPR@1%/TPR@5%/AUROC.
T
Clean
Parrot
DIPPER
PPL ↓
Samples / sent. ↓
5
0.71 / 0.83 / 0.95
0.32 / 0.56 / 0.81
0.27 / 0.40 / 0.74
4.61
1.3
6
0.98 / 1.00 / 1.00
0.94 / 0.99 / 1.00
0.93 / 0.98 / 0.99
4.61
2.2
7
0.98 / 1.00 / 1.00
0.98 / 0.99 / 1.00
0.98 / 0.99 / 1.00
5.33
6.4
Appendix
Table 10: Effect of the matching threshold T on C4 with Mistral-7B using the Global Bits detector. Detection results are reported as TPR@1% / TPR@5% / AUROC. Lower sampling cost is better.
Detection key
TPR@1% ↑
Correct key R⋆
0.99
Independent wrong key R′
0.08
Appendix
Table 11: Cross-key detection on C4 with Mistral-7B. Watermarked text is generated using R⋆ . TPR is measured at a target FPR of 1% .