Large language model watermarking embeds detectable statistical signals during decoding, but the resulting changes to token probabilities can degrade generation quality. This trade-off is particularly important for code, where small changes in token selection can break syntax or alter program behavior. Existing code watermarking methods mitigate this risk through entropy-based insertion or syntax-aware token selection, but they do not directly construct the watermark over the set of continuations admitted by the current grammar state. We propose Grammar-Guided Code Watermarking with Green Temperature (GTCW), which integrates grammar-constrained decoding with probability-aware watermarking. At each decoding step, GTCW restricts the candidate set to grammar-admissible tokens and partitions this support into keyed green and red subsets. At eligible high-entropy positions, green temperature reweights the green tokens according to the model's relative preferences, strengthening the watermark signal while retaining the grammar constraint. Across five models and five benchmarks spanning four programming languages, GTCW achieves a mean AUROC of 73.61%, compared with 67.83% for the strongest baseline, while maintaining a mean Pass@1 of 59.18% versus 59.58% for unwatermarked generation. Our implementation is available at https://github.com/hyundong98/GTCW .
Figures & tables
Figure 1: AUROC-Pass@1 tradeoff (left) and TPR@5%FPR (right), averaged equally over five models, five benchmarks, and three seeds. The star marks GTCW and the dashed line denotes unwatermarked Pass@1 in the left panel.
Figure 2: Overview of GTCW . (a) Grammar restricts the token support and contributes to keyed partitioning over admissible tokens. (b) At entropy-eligible positions, green temperature sharpens green logits around a shared anchor.
HumanEval+
MBPP+
HumanEvalPack
Model
Method
Python
Python
C++
Java
Go
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
Qwen2.5-C (3B)
Base
—
—
67.9
—
—
52.1
—
—
52.2
—
—
71.8
—
—
51.2
Unigram
68.1
18.9
65.0
82.4
38.5
43.5
75.5
28.5
46.8
67.2
16.9
63.4
71.2
15.5
45.7
KGW
68.1
22.0
63.4
81.2
39.7
47.6
76.7
33.9
45.3
65.5
14.4
66.9
68.6
22.0
47.6
SWEET
68.9
24.4
68.7
84.4
41.8
50.8
71.0
30.1
53.9
68.5
25.8
71.1
72.8
29.3
48.2
Table 1: Detection and functional correctness. AUC, T@5, and P@1 denote AUROC, TPR@5%FPR, and Pass@1, respectively. Values are mean percentages over three seeds. Gray shading indicates GTCW . Bold and underline mark the best and second-best watermarked methods, respectively, based on unrounded values.
HumanEval+
MBPP+
HumanEvalPack
AVG
Setting
Python
Python
C++
Java
Go
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
w/o Gram.
76.4
36.4
63.2
96.0
76.5
52.2
83.1
57.3
49.0
76.4
40.7
69.7
85.0
60.0
45.5
83.4
54.2
55.9
w/o Tg
68.7
27.0
69.7
86.8
50.4
50.0
72.3
35.4
49.4
67.8
21.3
68.9
76.7
29.9
50.0
74.5
32.8
57.6
GTCW
77.6
49.2
69.5
96.5
83.6
51.8
81.1
51.4
50.0
75.9
37.2
70.3
86.4
56.5
49.2
83.5
55.6
58.2
Table 2: Ablation of grammar guidance and green temperature on Qwen2.5-Coder (3B). All settings use the same entropy gate. AUC, T@5, and P@1 denote AUROC, TPR@5%FPR, and Pass@1. Values are mean percentages over three seeds.
HumanEval+
MBPP+
HumanEvalPack
Tg
Python
Python
C++
Java
Go
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
0.3
78.0
51.8
69.3
96.7
84.9
52.4
82.9
54.3
47.8
76.6
38.4
70.1
87.1
59.1
48.6
0.4
77.9
50.8
69.3
96.7
84.3
52.1
82.2
52.6
49.0
76.3
38.2
70.5
86.9
58.9
48.0
0.5
77.6
49.2
69.5
96.5
83.6
51.8
81.1
51.4
50.0
75.9
37.2
70.3
86.4
56.5
49.2
0.6
76.4
46.5
68.7
95.6
81.0
51.5
79.1
47.6
49.8
74.3
34.3
69.9
84.7
51.4
49.0
Table 3: Sensitivity to green temperature Tg on Qwen2.5-Coder (3B) with T0=1 . AUC, T@5, and P@1 denote AUROC, TPR@5%FPR, and Pass@1. Values are mean percentages over three seeds. Gray shading indicates the default Tg=0.5 .
Transformation
Setting
Unigram
KGW
SWEET
STONE
STA-1
GTCW
AUC
T@5
AUC
T@5
AUC
T@5
AUC
T@5
AUC
T@5
AUC
T@5
Clean
—
72.9
23.6
72.0
26.4
73.1
30.3
65.9
15.8
62.3
11.9
83.5
55.6
Renaming
25%
73.5
24.6
67.3
24.4
72.0
23.3
64.5
17.3
60.6
12.0
73.2
29.2
50%
74.3
26.1
66.9
23.7
71.3
22.7
63.7
17.8
59.5
12.0
72.4
27.8
75%
75.1
26.3
65.0
22.6
70.8
21.6
63.1
17.7
59.1
11.8
69.6
25.0
100%
75.4
26.4
64.0
22.4
70.0
21.3
62.9
17.5
58.1
11.2
69.4
24.6
Table 4: Detection after code transformations on Qwen2.5-Coder (3B). AUC and T@5 denote AUROC and TPR@5%FPR. Values average five benchmarks and three seeds. AVG averages the 13 transformation settings, excluding Clean. Gray shading indicates GTCW , with bold and underline marking the best and second-best methods.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
HumanEval+
MBPP+
HumanEvalPack
Model
Method
Python
Python
C++
Java
Go
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
AUC
T@5
P@1
Qwen2.5-C (3B)
Base
—
—
67.9
—
—
52.1
—
—
52.2
—
—
71.8
—
—
51.2
Unigram
68.1
18.9
65.0
82.4
38.5
43.5
75.5
28.5
46.8
67.2
16.9
63.4
71.2
15.5
45.7
KGW
68.1
22.0
63.4
81.2
39.7
47.6
76.7
33.9
45.3
65.5
14.4
66.9
68.6
22.0
47.6
SWEET
68.9
24.4
68.7
84.4
41.8
50.8
71.0
30.1
53.9
68.5
25.8
71.1
72.8
29.3
48.2
Appendix
Table 5: Full detection and functional-correctness results across five models and five benchmarks. AUC denotes AUROC, T@5 denotes TPR@5%FPR, and P@1 denotes Pass@1. Values are mean percentages over three seeds. Gray shading indicates GTCW . Bold and underline mark the best and second-best watermarked methods, respectively, based on unrounded values.
Method
Pass@1
Detected among
Correct and
Coverage
FPR
correct ( Q )
detected ( J )
Base
59.58
—
—
—
—
Unigram
55.00
10.12
5.29
100.00
4.84
KGW
55.40
12.51
6.70
99.97
4.85
SWEET
58.87
14.50
8.16
74.42
3.74
STONE
57.03
9.84
5.38
99.90
4.83
Appendix
Table 6: Detection on functionally correct outputs at TPR@5%FPR. Q denotes detection among correct outputs, and J denotes outputs that are both correct and detected. Coverage is the fraction with finite scores. Values are percentages averaged over five models, five benchmarks, and three seeds.
Method
Seconds/completion
Tokens/completion
Tokens/second
Base
4.28
175.85
64.88
Unigram
4.62
192.68
64.93
KGW
4.56
187.60
63.70
SWEET
4.41
179.10
63.37
STONE
4.46
180.10
63.07
STA-1
4.40
178.45
63.49
Appendix
Table 7: Generation wall-clock time, output length, and throughput. Values are averaged equally over five models, five benchmarks, and three seeds. Throughput is averaged per completion.
Figure 3: Functionally correct and detected outputs for MBPP+/796 using Qwen2.5-Coder (3B). All displayed completions pass the augmented tests and exceed their corresponding TPR@5%FPR detection thresholds. Z denotes the detector score and G/N the number of green tokens among scored positions. Green highlighting marks scored green tokens.