Leading multi-bit watermarking methods for language models encode messages by biasing the model's next-token probabilities, creating a trade-off between message recovery and text quality. Their decoders typically return the highest-scoring candidate from accumulated token-level evidence, without a certified abstention rule that bounds the probability of outputting an incorrect message. We introduce CertMark, a distribution-preserving multi-bit watermark with certified decoding. Rather than modifying probabilities, CertMark uses the embedded message to seed an exact Gumbel-max sampler, thereby preserving the model's original sampling distribution. We propose two scalable decoders: a model-agnostic, text-only decoder and a model-aware variant that leverages the original next-token distributions for stronger recovery. Both support certified abstention with mathematical bounds on the probability of returning an incorrect message. Across text completion, summarization, and story generation, CertMark matches the perplexity of unwatermarked text while reliably recovering multi-bit messages. The model-aware decoder further achieves higher bit accuracy than probability-biasing baselines. Our code is publicly available at https://github.com/Batorskq/CertMark.
Figures & tables
Model and text
Encoder
V
vocabulary
h
context length in tokens
x
prompt
ct=yt−h:t−1
context of step t
y1:n , yt , y<t
text, its t -th token, its prefix
it
chunk carried at step t
Y1:n
emitted text as a random variable
Ct
contexts used so far
pt(⋅)=p(⋅∣x,y<t)
deployed sampler’s law at step t
at
argument tuple of step t
Key and randomness
ut(v)
uniform attached to token v at step t
Table 1: Notation used in the method.
Method
T=150
T=200
T=250
T=300
Avg.
Distortion (avg. over T )
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
Top-1 ↑
Top-5 ↑
R-1 ↑
R-L ↑
Text Completion
CycleShift
100.00
7.97
100.00
7.58
100.00
7.05
100.00
7.44
100.00
7.51
55.02
84.96
0.290
0.170
DepthW
95.75
7.40
100.00
6.96
100.00
6.52
97.25
6.68
98.25
6.89
57.39
85.87
0.290
0.168
StealthInk
89.75
6.15
90.50
5.79
95.25
5.73
94.25
5.61
92.44
5.82
62.81
89.16
0.308
0.179
MPAC
97.25
7.70
98.00
7.52
98.75
7.03
99.00
7.09
98.25
7.33
55.60
85.34
0.293
0.170
Table 2: Main benchmark on Qwen3.5-4B: C4 text completion, CNN/DailyMail summarization and WritingPrompts story generation with L=8 and token budgets T∈{150,200,250,300} . Perplexity is measured under Qwen3.5-9B; the distortion columns are defined in Section 4 . The four CertMark rows are the model-agnostic and the model-aware decoder, each with one 8-bit chunk and with 2-bit chunks. Colours mark the best, second-best and third-best method in each column.
Qwen3.5-4B
Llama-3.1-8B
Method
BA ↑
PPL ↓
BLEU ↑
BA ↑
PPL ↓
BLEU ↑
RSBH
59.38
2.51
20.64
57.62
3.49
24.84
MPAC
70.00
2.58
20.98
64.62
3.50
26.80
XMark
74.38
2.49
20.32
71.25
3.46
25.99
StealthInk
64.88
2.22
23.00
55.75
3.19
27.08
CertMark model-agnostic 16-bit
61.50
2.18
23.15
48.88
3.08
26.34
Table 3: WMT14 German-to-English translation.
Qwen3.5-4B
Llama-3.1-8B
Method
BA ↑
PPL ↓
BLEU ↑
BA ↑
PPL ↓
BLEU ↑
RSBH
59.38
2.51
20.64
57.62
3.49
24.84
MPAC
70.00
2.58
20.98
64.62
3.50
26.80
XMark
74.38
2.49
20.32
71.25
3.46
25.99
StealthInk
64.88
2.22
23.00
55.75
3.19
27.08
CertMark model-agnostic 16-bit
61.50
2.18
23.15
48.88
3.08
26.34
Table 3: WMT14 German-to-English translation.
Qwen3.5-4B
Llama-3.1-8B
Method
BA ↑
PPL ↓
BSc. ↑
BA ↑
PPL ↓
BSc. ↑
RSBH
85.03
5.38
.8208
89.25
4.63
.8341
MPAC
87.28
5.46
.8170
88.38
4.94
.8279
XMark
93.38
5.48
.8199
96.00
4.74
.8351
StealthInk
79.81
4.52
.8222
80.69
3.75
.8392
CertMark model-agnostic 16-bit
99.75
4.29
.8262
99.44
3.66
.8365
Table 4: Long C4 completion ( L=64 ).
CertMark model-agnostic
CertMark model-aware
δ
10−4
10−3
0.01
0.05
0.1
0.2
0.5
10−4
10−3
0.01
0.05
0.1
0.2
0.5
Qwen3.5-4B
cert. error ( ≤δ )
.0000
.0000
.0063
.0221
.0305
.0537
.1168
.0000
.0011
.0074
.0147
.0168
.0274
.0474
null cert. rate ( ≈δ )
.0021
.0063
.0137
.0621
.1200
.2147
.4074
.0000
.0000
.0074
.0442
.1000
.1895
.3863
abst., completion
.178
.129
.082
.056
.044
.027
.013
.040
.022
.009
.004
.004
.002
.000
abst., summarization
.902
.840
.727
.627
.582
.498
.324
.604
.509
.387
.251
.207
.158
.069
Table 5: Empirical check of the certificate (Theorem 4 , and its model-aware form in Appendix A ). Each column fixes a level δ : the decoder returns a chunk’s message only if δ^c≤δ and abstains ( ⊥ ) otherwise. Both decoders read the same saved generations of the message-length sweep ( T=150 ; L∈{8,…,32} on completion and summarization, L=8 on story generation; 50 users each), 950 watermarked chunks and 950 paired unwatermarked chunks per model. Every entry is a fraction of chunks. Cert. error : watermarked chunks on which a wrong message is returned; the theorem bounds its probability by δ , and with 950 chunks a single error is .0011 . Null cert. rate : unwatermarked chunks that are nevertheless certified, which should stay near δ . Abst. : watermarked chunks of each task on which the decoder abstains; lower means more messages recovered with a certificate. The model-aware decoder uses the robust score with ε=0.1 . For every row, lower is better.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen3.5-4B
Llama-3.1-8B
Method
Gen.
Dec.
Gen.
Dec.
CycleShift
19.6
.15
15.0
.07
DepthW
19.0
11.56
14.9
12.03
StealthInk
20.0
.18
15.3
.10
MPAC
18.9
.03
14.8
.03
RSBH
19.7
2.18
15.1
1.18
Appendix
Table 6: Runtime on one H100. Gen.: ms/token; Dec.: s/user.
Method
T=150
T=200
T=250
T=300
Avg.
Distortion (avg. over T )
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
Top-1 ↑
Top-5 ↑
R-1 ↑
R-L ↑
Text Completion
CycleShift
100.00
5.63
100.00
5.75
100.00
5.59
100.00
5.51
100.00
5.62
57.62
88.55
0.321
0.186
DepthW
98.25
5.28
99.50
5.12
99.00
5.12
100.00
5.19
99.19
5.18
60.75
89.56
0.337
0.195
StealthInk
90.00
4.61
91.00
4.35
94.25
4.58
96.25
4.28
92.88
4.46
64.98
92.37
0.347
0.202
MPAC
96.75
5.84
98.00
5.48
99.25
5.45
99.00
5.67
98.25
5.61
57.66
88.34
0.329
0.188
Appendix
Table 7: Table 2 repeated with Llama-3.1-8B (completion) and Llama-3.1-8B-Instruct (chat tasks) as the generating model, with perplexity under the other 8B sibling of the family since the 70B oracle does not fit one GPU. Every number is measured in our harness; layout and colours follow the main table.
Method
T=150
T=200
T=250
T=300
Avg.
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
BA ↑
PPL ↓
Text Completion
StealthInk
89.75
6.15
90.50
5.79
95.25
5.73
94.25
5.61
92.44
5.82
BiMark
98.25
5.79
99.75
5.81
99.75
5.40
99.25
5.26
99.25
5.56
MirrorMark
100.00
5.83
100.00
5.66
100.00
5.12
100.00
5.27
100.00
5.47
CertMark model-agnostic 8-bit
100.00
5.65
100.00
5.23
100.00
5.29
100.00
5.11
100.00
5.32
Appendix
Table 8: Distortion-free multi-bit schemes on Qwen3.5-4B: StealthInk, BiMark, MirrorMark and the four CertMark rows of Table 2 , under the same protocol, tasks and budgets, with perplexity under Qwen3.5-9B; every number is measured in our harness. Red marks the best value in each column.
Department of AI, DGIST, Daegu, Republic of Korea · Department of ICT Convergence, University of Ulsan, Ulsan, Republic of Korea · Department of EECS, DGIST, Daegu, Republic of Korea