TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
Abstract
Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked. A fixed green list needs no such context, but it favors the same tokens at every position, so these tokens appear more often in watermarked text. An attacker who compares token frequencies in watermarked and unwatermarked text can recover the list and forge text that the provider's own detector accepts. We present TANGO, a watermark for masked-diffusion language models that keys each new token to a nearby token that is already unmasked. A secret key splits the vocabulary into color classes, and TANGO biases the new token toward a color determined by the key and the nearby token's color. The watermark is therefore embedded in pairs of tokens. Because the favored color changes from position to position, token frequencies stay much closer to those of unwatermarked text than under a fixed green list. Detection needs only the text and the key, and it does not assume any unmasking order. On two masked-diffusion models, TANGO detects nearly all unedited watermarked texts and most edited ones, and frequency attacks that forge the fixed green list fail against it.
Figures & tables
| TPR@1%FPR | |||||||
| model | method | AUROC | clean | del30 | syn30 | ins20 | PPL |
| LLaDA | TANGO | 0.998 | 7.2 | ||||
| LLaDA | Red–green list | 0.993 | 7.6 | ||||
| LLaDA | Gumbel | 0.815 | 4.6 | ||||
| LLaDA | unwatermarked | – | – | – | – | – | 4.2 |
| Dream | TANGO | 1.000 | 7.8 | ||||
| TPR@1%FPR | ||||||||
| axis | setting | AUROC | clean | del30 | syn30 | ins20 | PPL | null FPR (%) |
| colors | 0.989 | 0.92 | 0.62 | 0.71 | 0.83 | 20.4 | 0.0 | |
| 1.000 | 1.00 | 0.92 | 0.96 | 1.00 | 22.2 | 0.0 | ||
| 1.000 | 1.00 | 0.79 | 0.83 | 0.96 | 26.8 | 0.5 | ||
| 1.000 | 1.00 | 0.83 | 0.83 | 1.00 | 29.9 | 0.5 | ||
| 0.998 | 0.96 | 0.88 | 0.96 | 0.96 | 29.5 | 0.0 | ||
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| result | claim in words | assumes | appendix |
|---|---|---|---|
| Proposition 1 | Unwatermarked text gives a count. | uniform colors, (I) | A.2 |
| Proposition 3 | A uniformly random balanced coloring gives every class mass close to unless the context is peaked. | hash coloring | A.2 |
| Proposition 2 | The expected score is . | definitions only, fixed | A.3 |
| Corollary 1 | The expected score reaches once . | , constant | A.3 |
| Lemma 1 † | Unenforced pairs match at the chance rate, so . | (S), (U ′ ), (B), single tap, | A.3 |
| Corollary 2 | The false-negative rate decays exponentially in . | (I) | A.4 |
| symbol | meaning |
|---|---|
| , | vocabulary and its size, |
| tokens of a text of length | |
| state after denoising steps, the partly unmasked sequence | |
| , | number of colors, and with arithmetic modulo |
| balanced coloring fixed by the key; class has tokens | |
| , , | set of tap lags, one lag, and its coefficient in |
| coloring | null source | mean | std | KS | FPR |
|---|---|---|---|---|---|
| unwhitened | natural | 1.10 | 0.403 | 1.00% | |
| unwhitened | random tokens | 1.04 | 0.047 | 0.00% | |
| ZCA | natural | 1.17 | 0.274 | 0.33% | |
| ZCA | random tokens | 1.00 | 0.057 | 0.00% |
| text | lag | class mass ( ) | null rate ( ) | null mean ( ) |
|---|---|---|---|---|
| natural | 1 | |||
| natural | 2 | |||
| model | 1 | |||
| model | 2 |
| residue | AUROC | clean | mean | null mean | del30 | syn30 | ins20 | PPL GPT-2 |
|---|---|---|---|---|---|---|---|---|
| 1.000 | 1.00 | 6.33 | 0.79 | 0.92 | 0.93 | 27.6 | ||
| 0.999 | 0.97 | 6.59 | 0.32 | 0.58 | 0.68 | 26.1 | ||
| 1.000 | 1.00 | 6.80 | 0.76 | 0.98 | 0.99 | 25.2 | ||
| tap-keyed | 0.991 | 0.97 | 6.02 | 0.91 | 0.93 | 0.97 | 26.9 |
| method | frequency shift (TV) | key recovery AUC | forgery success |
|---|---|---|---|
| TANGO | 0.217 | 0.504 | 0.00 |
| Red–green list | 0.340 | 0.817 | 1.00 |
| model | method | ||||
|---|---|---|---|---|---|
| LLaDA | TANGO | 0.62 / 0.12 | 0.62 / 0.18 | 0.62 / 0.48 | 0.62 / 0.74 |
| LLaDA | Red–green list | 0.63 / 0.51 | 0.64 / 0.74 | 0.65 / 0.89 | 0.66 / 0.96 |
| LLaDA | Gumbel | 0.53 / 0.17 | 0.54 / 0.25 | 0.53 / 0.24 | 0.53 / 0.56 |
| LLaDA | DLM watermark, defaults | 0.54 / – | 0.55 / – | 0.54 / – | 0.54 / – |
| LLaDA | DLM watermark, matched, bias | 0.53 / – | 0.53 / – | 0.55 / – | 0.55 / – |
| LLaDA | DLM watermark, matched, bias | 0.55 / – | 0.55 / – | 0.56 / – | 0.57 / – |
| detector | mean | success (1% FPR) | success ( ) | PPL GPT-2 | |
|---|---|---|---|---|---|
| 4 | TANGO | 0.08 | 0.04 | 0.00 | 13.5 |
| 4 | Red–green list | 4.15 | 0.51 | 0.56 | 15.2 |
| 8 | TANGO | 2.62 | 0.55 | 0.23 | 21.1 |
| 8 | Red–green list | 10.56 | 0.98 | 0.98 | 30.2 |
| TPR@1%FPR | PPL (control) | |||||||||
| method | AUROC | clean | del30 | syn30 | ins20 | Qwen | GPT-2 | freq. AUC | pair AUC | forge |
| TANGO ( , either side) | 0.998 | 0.97 | 0.68 | 0.92 | 0.91 | 7.0 (4.1) | 18.9 (12.9) | 0.50 | 0.62 | 0.74 |
| Red–green list, bias | 0.993 | 0.98 | 0.96 | 0.94 | 0.98 | 7.4 (4.1) | 17.6 (12.9) | 0.82 | 0.66 | 0.96 |
| Gumbel rule | 0.815 | 0.15 | 0.10 | 0.07 | 0.07 | 4.5 (4.1) | 13.6 (12.9) | – | 0.53 | 0.56 |
| DLM watermark, defaults (bias ) | 0.843 | 0.37 | 0.30 | 0.22 | 0.24 | 6.5 (6.1) | 39.9 (71.1) | – | 0.54 | – |
| DLM watermark, matched, bias | 0.808 | 0.19 | 0.16 | 0.09 | 0.09 | 4.2 (4.0) | 24.1 (21.5) | – | 0.55 | – |
| setting | AUROC | clean | del30 | syn30 | ins20 | PPL GPT-2 | PPL Qwen |
|---|---|---|---|---|---|---|---|
| TANGO | 0.998 | 0.97 | 0.61 | 0.84 | 0.85 | 16.8 | 6.1 |
| TANGO ( Table 1 ) | 0.998 | 0.97 | 0.68 | 0.92 | 0.91 | 19.4 | 7.2 |
| TANGO | 0.999 | 0.99 | 0.81 | 0.95 | 0.96 | 21.3 | 8.4 |
| TANGO | 1.000 | 1.00 | 0.93 | 0.99 | 0.98 | 31.4 | 13.5 |
| TANGO | 1.000 | 1.00 | 0.97 | 0.98 | 0.99 | 51.8 | 23.4 |
| Red–green list, bias | 0.940 | 0.58 | 0.38 | 0.39 | 0.69 | 13.5 | 4.6 |
| TPR@1%FPR | |||||||
|---|---|---|---|---|---|---|---|
| method | AUROC | clean | del30 | syn30 | ins20 | PPL GPT-2 | degenerate |
| TANGO ( , either side) | 1.000 | 1.00 [1.00,1.00] | 0.73 [0.65,0.81] | 0.98 [0.95,1.00] | 0.97 [0.93,1.00] | 15.2 | 31% |
| Red–green list | 1.000 | 1.00 [1.00,1.00] | 0.98 [0.95,1.00] | 0.98 [0.95,1.00] | 0.98 [0.95,1.00] | 18.1 | 62% |
| unwatermarked | – | – | – | – | – | 9.9 | 20% |
| key-aware | random | |||||
|---|---|---|---|---|---|---|
| budget | mean | PPL GPT-2 | det. rate | mean | PPL GPT-2 | det. rate |
| 10% | 4.34 | 28.1 | 0.62 | 4.42 | 28.5 | 0.60 |
| 25% | 3.11 | 44.3 | 0.27 | 2.81 | 44.2 | 0.20 |
| 50% | 1.54 | 76.8 | 0.02 | 1.03 | 95.2 | 0.02 |
| attack | TPR@1%FPR (95% CI) | per-seed std |
|---|---|---|
| clean | 0.98 [0.95,1.00] | 0.04 |
| del10 | 0.92 [0.85,0.97] | 0.08 |
| del30 | 0.62 [0.52,0.72] | 0.18 |
| syn30 | 0.85 [0.78,0.92] | 0.11 |
| syn50 | 0.61 [0.51,0.71] | 0.19 |
| ins20 | 0.91 [0.84,0.96] | 0.09 |
| quantity | mean | std |
|---|---|---|
| clean AUROC | 0.998 | 0.002 |
| mean on natural text | 0.28 |
| coloring | AUROC | clean | del10 | del30 | syn30 | syn50 | ins20 | PPL GPT-2 | null FPR (%) |
|---|---|---|---|---|---|---|---|---|---|
| quantile, | 0.989 | 0.92 | 0.79 | 0.62 | 0.71 | 0.58 | 0.83 | 20.4 | 0.0 |
| quantile, | 1.000 | 1.00 | 1.00 | 0.92 | 0.96 | 0.75 | 1.00 | 22.2 | 0.0 |
| quantile, | 1.000 | 1.00 | 0.96 | 0.79 | 0.83 | 0.67 | 0.96 | 26.8 | 0.5 |
| quantile, | 1.000 | 1.00 | 1.00 | 0.83 | 0.83 | 0.62 | 1.00 | 29.9 | 0.5 |
| quantile, | 0.998 | 0.96 | 0.96 | 0.88 | 0.96 | 0.92 | 0.96 | 29.5 | 0.0 |
| cluster, | 0.817 | 0.38 | 0.25 | 0.12 | 0.17 | 0.08 | 0.46 | 13.3 | 60.0 |
| taps | AUROC | clean | del10 | del30 | syn30 | syn50 | ins20 | PPL GPT-2 | null FPR (%) |
|---|---|---|---|---|---|---|---|---|---|
| 0.990 | 0.96 | 0.96 | 0.92 | 0.83 | 0.71 | 0.96 | 25.6 | 0.0 | |
| 1.000 | 1.00 | 1.00 | 0.92 | 0.96 | 0.92 | 0.96 | 25.0 | 0.0 | |
| 0.990 | 0.96 | 0.79 | 0.25 | 0.58 | 0.17 | 0.79 | 19.6 | 0.0 | |
| 1.000 | 1.00 | 0.88 | 0.50 | 0.62 | 0.62 | 0.75 | 18.5 | 0.0 | |
| 0.986 | 0.83 | 0.62 | 0.25 | 0.25 | 0.04 | 0.54 | 16.7 | 0.0 | |
| 0.984 | 0.83 | 0.50 | 0.12 | 0.38 | 0.25 | 0.42 | 18.9 | 0.0 |
| setting | AUROC | clean | del30 | syn30 | ins20 | enforced fraction | PPL GPT-2 | PPL Qwen | pair AUC / forge | |
|---|---|---|---|---|---|---|---|---|---|---|
| one-sided, (reference) | 0.999 | 0.98 | 0.79 | 0.98 | 0.98 | 0.61 | 24.0 | 9.7 | 0.61 / 0.36 | |
| one-sided, | 0.999 | 0.98 | 0.75 | 0.79 | 0.96 | 0.62 | 19.6 | 7.4 | 0.61 / 0.26 | |
| either side, | 1.000 | 1.00 | 0.94 | 1.00 | 1.00 | 0.74 | 31.1 | 12.7 | 0.62 / 0.47 | |
| either side, | 1.000 | 1.00 | 0.77 | 0.96 | 0.96 | 0.75 | 19.6 | 7.8 | 0.63 / 0.38 | |
| either side, | 1.000 | 1.00 | 0.79 | 0.92 | 0.98 | 0.76 | 17.8 | 7.0 | 0.62 / 0.40 | |
| either side, | 0.998 | 0.96 | 0.73 | 0.81 | 0.81 | 0.77 | 16.0 | 5.9 | 0.61 / 0.23 |