Large language model watermarking embeds detectable statistical signals during decoding, but the resulting changes to token probabilities can degrade generation quality. This trade-off is particularly important for code, where small changes in token selection can break syntax or alter program behavior. Existing code watermarking methods mitigate this risk through entropy-based insertion or syntax-aware token selection, but they do not directly construct the watermark over the set of continuations admitted by the current grammar state. We propose Grammar-Guided Code Watermarking with Green Temperature (GTCW), which integrates grammar-constrained decoding with probability-aware watermarking. At each decoding step, GTCW restricts the candidate set to grammar-admissible tokens and partitions this support into keyed green and red subsets. At eligible high-entropy positions, green temperature reweights the green tokens according to the model's relative preferences, strengthening the watermark signal while retaining the grammar constraint. Across five models and five benchmarks spanning four programming languages, GTCW achieves a mean AUROC of 73.61%, compared with 67.83% for the strongest baseline, while maintaining a mean Pass@1 of 59.18% versus 59.58% for unwatermarked generation. Our implementation is available at https://github.com/hyundong98/GTCW .
Figures & tables
Figure 1: AUROC-Pass@1 tradeoff (left) and TPR@5%FPR (right), averaged equally over five models, five benchmarks, and three seeds. The star marks GTCW and the dashed line denotes unwatermarked Pass@1 in the left panel.
Figure 2: Overview of GTCW . (a) Grammar restricts the token support and contributes to keyed partitioning over admissible tokens. (b) At entropy-eligible positions, green temperature sharpens green logits around a shared anchor.
HumanEval+
MBPP+
HumanEvalPack
Model
Method
Python
Python
C++
Java
Go
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
Qwen2.5-C (3B)
Base
—
—
67.9
—
—
52.1
—
—
52.2
—
—
71.8
—
—
51.2
Unigram
68.1
18.9
65.0
82.4
38.5
43.5
75.5
28.5
46.8
67.2
16.9
63.4
71.2
15.5
45.7
KGW
68.1
22.0
63.4
81.2
39.7
47.6
76.7
33.9
45.3
65.5
14.4
66.9
68.6
22.0
47.6
SWEET
68.9
24.4
68.7
84.4
41.8
50.8
71.0
30.1
53.9
68.5
25.8
71.1
72.8
29.3
48.2
Table 1: Detection and functional correctness. AUC, T@5, and P@1 denote AUROC, TPR@5%FPR, and Pass@1, respectively. Values are mean percentages over three seeds. Gray shading indicates GTCW . Bold and underline mark the best and second-best watermarked methods, respectively, based on unrounded values.
HumanEval+
MBPP+
HumanEvalPack
AVG
Setting
Python
Python
C++
Java
Go
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
w/o Gram.
76.4
36.4
63.2
96.0
76.5
52.2
83.1
57.3
49.0
76.4
40.7
69.7
85.0
60.0
45.5
83.4
54.2
55.9
w/o Tg
68.7
27.0
69.7
86.8
50.4
50.0
72.3
35.4
49.4
67.8
21.3
68.9
76.7
29.9
50.0
74.5
32.8
57.6
GTCW
77.6
49.2
69.5
96.5
83.6
51.8
81.1
51.4
50.0
75.9
37.2
70.3
86.4
56.5
49.2
83.5
55.6
58.2
Table 2: Ablation of grammar guidance and green temperature on Qwen2.5-Coder (3B). All settings use the same entropy gate. AUC, T@5, and P@1 denote AUROC, TPR@5%FPR, and Pass@1. Values are mean percentages over three seeds.
HumanEval+
MBPP+
HumanEvalPack
Tg
Python
Python
C++
Java
Go
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
0.3
78.0
51.8
69.3
96.7
84.9
52.4
82.9
54.3
47.8
76.6
38.4
70.1
87.1
59.1
48.6
0.4
77.9
50.8
69.3
96.7
84.3
52.1
82.2
52.6
49.0
76.3
38.2
70.5
86.9
58.9
48.0
0.5
77.6
49.2
69.5
96.5
83.6
51.8
81.1
51.4
50.0
75.9
37.2
70.3
86.4
56.5
49.2
0.6
76.4
46.5
68.7
95.6
81.0
51.5
79.1
47.6
49.8
74.3
34.3
69.9
84.7
51.4
49.0
Table 3: Sensitivity to green temperature Tg on Qwen2.5-Coder (3B) with T0=1 . AUC, T@5, and P@1 denote AUROC, TPR@5%FPR, and Pass@1. Values are mean percentages over three seeds. Gray shading indicates the default Tg=0.5 .
Transformation
Setting
Unigram
KGW
SWEET
STONE
STA-1
GTCW
AUC
T@5
AUC
T@5
AUC
T@5
AUC
T@5
AUC
T@5
AUC
T@5
Clean
—
72.9
23.6
72.0
26.4
73.1
30.3
65.9
15.8
62.3
11.9
83.5
55.6
Renaming
25%
73.5
24.6
67.3
24.4
72.0
23.3
64.5
17.3
60.6
12.0
73.2
29.2
50%
74.3
26.1
66.9
23.7
71.3
22.7
63.7
17.8
59.5
12.0
72.4
27.8
75%
75.1
26.3
65.0
22.6
70.8
21.6
63.1
17.7
59.1
11.8
69.6
25.0
100%
75.4
26.4
64.0
22.4
70.0
21.3
62.9
17.5
58.1
11.2
69.4
24.6
Table 4: Detection after code transformations on Qwen2.5-Coder (3B). AUC and T@5 denote AUROC and TPR@5%FPR. Values average five benchmarks and three seeds. AVG averages the 13 transformation settings, excluding Clean. Gray shading indicates GTCW , with bold and underline marking the best and second-best methods.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
HumanEval+
MBPP+
HumanEvalPack
Model
Method
Python
Python
C++
Java
Go
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
Qwen2.5-C (3B)
Base
—
—
67.9
—
—
52.1
—
—
52.2
—
—
71.8
—
—
51.2
Unigram
68.1
18.9
65.0
82.4
38.5
43.5
75.5
28.5
46.8
67.2
16.9
63.4
71.2
15.5
45.7
KGW
68.1
22.0
63.4
81.2
39.7
47.6
76.7
33.9
45.3
65.5
14.4
66.9
68.6
22.0
47.6
SWEET
68.9
24.4
68.7
84.4
41.8
50.8
71.0
30.1
53.9
68.5
25.8
71.1
72.8
29.3
48.2
Appendix
Table 5: Full detection and functional-correctness results across five models and five benchmarks. AUC denotes AUROC, T@5 denotes TPR@5%FPR, and P@1 denotes Pass@1. Values are mean percentages over three seeds. Gray shading indicates GTCW . Bold and underline mark the best and second-best watermarked methods, respectively, based on unrounded values.
Method
Pass@1
Detected among
Correct and
Coverage
FPR
correct ( Q )
detected ( J )
Base
59.58
—
—
—
—
Unigram
55.00
10.12
5.29
100.00
4.84
KGW
55.40
12.51
6.70
99.97
4.85
SWEET
58.87
14.50
8.16
74.42
3.74
STONE
57.03
9.84
5.38
99.90
4.83
Appendix
Table 6: Detection on functionally correct outputs at TPR@5%FPR. Q denotes detection among correct outputs, and J denotes outputs that are both correct and detected. Coverage is the fraction with finite scores. Values are percentages averaged over five models, five benchmarks, and three seeds.
Method
Seconds/completion
Tokens/completion
Tokens/second
Base
4.28
175.85
64.88
Unigram
4.62
192.68
64.93
KGW
4.56
187.60
63.70
SWEET
4.41
179.10
63.37
STONE
4.46
180.10
63.07
STA-1
4.40
178.45
63.49
Appendix
Table 7: Generation wall-clock time, output length, and throughput. Values are averaged equally over five models, five benchmarks, and three seeds. Throughput is averaged per completion.
Figure 3: Functionally correct and detected outputs for MBPP+/796 using Qwen2.5-Coder (3B). All displayed completions pass the augmented tests and exceed their corresponding TPR@5%FPR detection thresholds. Z denotes the detector score and G/N the number of green tokens among scored positions. Green highlighting marks scored green tokens.
With the rapid development of Large Language Models (LLMs), text watermarking has emerged as a crucial technique for identifying machine-generated content. However, directly applying existing logits-based watermarking methods to code generation remains challenging, since the low-entropy nature of code exacerbates the trade-off between code quality and watermark detectability. In this paper, we propose a novel code watermarking approach called Grammar-Driven Watermark (GDW) for LLMs. GDW preserves syntactic validity through a grammar-guided three-level masking mechanism and injects watermark signals via structural role-aware modulation, assigning a stronger bias to content-bearing tokens while applying a more conservative bias to syntax-critical tokens. Aligning with the generation process, we further design a role-aware weighted detection statistic to improve detectability. Experiments across multiple programming languages, models, and decoding strategies show that GDW establishes a stronger quality-detectability trade-off frontier than existing methods, while maintaining robustness against variable-renaming attacks.
Attributing code to the large language model that produced it is essential for provenance, licensing, and misuse accountability, yet no deployed watermark meets this need. Generation-time schemes require access to the producing model and cannot be applied to third-party code, while post-hoc schemes work on any code but carry at most 4 bits of payload, far too few to distinguish the many deployed model configurations. We present multi-channel spread-spectrum watermarking, the first post-hoc, training-free code watermark with a 24-bit payload and formal robustness guarantees. The scheme encodes bits in variable naming conventions and in eight pairs of semantically equivalent code patterns, and a keyed pseudo-random permutation maps every site to a codeword bit so that each bit receives multiple independent votes. Majority voting absorbs distributed corruption, while an outer Reed-Solomon code recovers the identifier when concentrated channel attacks defeat the vote, yielding provable robustness bounds for formatting, syntactic, and structural attacks. Across 1,750 Python files from CodeNet and from GPT-4.1 and Llama-4 generations, the watermark achieves 100% clean-detection accuracy with zero false positives. Under 17 attack types, it recovers the identifier at 97.6% accuracy under 8 variable renames and 94.1% under 10% random per-site corruption, while the strongest post-hoc baseline collapses to 0% under any single-transform attack. Embedding and detection together take under 200 ms on CPU without training data or GPU.
Soohyeon Choi, Debin Gao, Yue Duan
Singapore Management University Singapore Singapore
Watermarking has emerged as a promising technique for tracing the authorship of content generated by large language models (LLMs). Among existing approaches, the KGW scheme is particularly attractive due to its versatility, efficiency, and effectiveness in natural language generation. However, KGW's effectiveness degrades significantly under low-entropy settings such as code generation and mathematical reasoning. A crucial step in the KGW method is random vocabulary partitioning, which enables adjustments to token selection based on specific preferences. Our study revealed that the next-token probability distribution plays an critical role in determining how much, or even whether, we can modify token selection and, consequently, the effectiveness of watermarking. We refer to this characteristic, associated with the probability distribution of each token prediction, as \emph{watermark strength.} In cases of random vocabulary partitioning, the lower bound of watermark strength is dictated by the next-token probability distribution. However, we found that, by redesigning the vocabulary partitioning algorithm, we can potentially raise this lower bound. In this paper, we propose SSG (\textbf{S}ort-then-\textbf{S}plit by \textbf{G}roups), a method that partitions the vocabulary into two logit-balanced subsets. This design lifts the lower bound of watermark strength for each token prediction, thereby improving watermark detectability. Experiments on code generation and mathematical reasoning datasets demonstrate the effectiveness of SSG.