Could LLM Watermark Detection be Public?
Organizations: FAIR, Meta Superintelligence Labs · University of Maryland, College Park
Abstract
Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback. However, watermarks are already vulnerable to uninformed tampering attacks. We thus first quantify whether a public detector would be an additional liability in a deployment setting at varying levels of access, from token-level scores to a binary verdict. Second, we introduce a split-key public-private watermarking method that exposes one key through a public detector while keeping the other for full verification and forensics. An informed attacker can only move the public signal, creating an imbalance between public and private scores. We introduce a statistical test for this imbalance, and combine it with the full key verdict in a two-stage mechanism. Third, we evaluate the split-key method on a wide range of removal and forgery attacks, comparing the uninformed to detector-informed settings. Public detection improves removal only at small edit budgets, since plain rephrasing already strips the watermark at a lower quality cost, but it does enable forgery, which the private pipeline can identify. Overall, releasing half of the watermark enables transparency and interoperability, and tampering with the released half stays detectable. This bounds the provider's liability and questions the need to keep detectors fully private.
Figures & tables
| Removal | Forgery | |||||||||||
| Rephrase (Add.) | Word edits (10%) | Rephrase (Add.) | Word edits (10%) | |||||||||
| Detector: | no-box | black-box | white-box | no-box | black-box | white-box | no-box | black-box | white-box | no-box | black-box | white-box |
| Gumbel-max | 70.4% | 81.4% | 99.9% | 1.5% | 1.6% | 72.2% | 0.0% | 0.6% | 20.8% | 0.2% | 0.3% | 56.3% |
| Gumbel-max | 82.0% | 92.9% | 99.5% | 6.7% | 8.4% | 97.8% | 0.1% | 0.3% | 6.2% | 0.1% | 0.1% | 48.3% |
| Maryland | 90.1% | 96.1% | 100.0% | 8.6% | 10.7% | 15.9% | 0.4% | 1.1% | 78.6% | 0.1% | 0.5% | 1.3% |
| Maryland | 95.2% | 98.9% | 100.0% | 34.4% | 40.4% | 98.7% | 0.1% | 0.9% | 38.8% | 0.1% | 0.3% | 1.9% |
| Removal | Forgery | |||||||||||||
| Rephrase | Word edits | Rephrase | Word edits | |||||||||||
| Metric | No-box | Adap. ( ) | Adap. ( ) | Adap. ( ) | No-box | Adap. ( ) | Adap. ( ) | Adap. ( ) | Adap. ( ) | Adap. ( ) | Adap. ( ) | Adap. ( ) | Adap. ( ) | Adap. ( ) |
| ASR pub | 92.3% | 99.5% | 99.9% | 100.0% | 62.5% | 99.9% | 99.8% | 99.8% | 18.1% | 19.6% | 20.8% | 99.8% | 99.9% | 99.9% |
| ASR fus | 82.0% | 88.5% | 95.1% | 99.4% | 26.9% | 8.8% | 30.8% | 99.3% | 4.8% | 7.5% | 5.5% | 45.8% | 69.8% | 85.1% |
| ADR | 0.4% | 0.2% | 0.2% | 0.7% | 0.0% | 0.0% | 0.3% | 99.4% | 0.0% | 1.3% | 5.5% | 10.9% | 36.0% | 66.2% |
| ASR net | 80.7% | 88.1% | 94.9% | 98.7% | 25.6% | 8.8% | 30.7% | 0.6% | 3.2% | 5.9% | 4.0% | 40.8% | 44.7% | 28.8% |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Watermark test | Tampering test | Either stage | ||
| Detection | Tampering | ||||
| Permutations | – | ||||
| Resolvable | analytic | ||||
| Test duration (ms) | 43.9 | 23.3 | 276.9 | 1105.1 | 5518.5 |
| Attack | Binomial | Permutation |
| Untampered (FPR) | 0.3 | 0.1 |
| Rephrase removal | 86.9 | 82.4 |
| Rephrase removal (additive) | 40.1 | 25.1 |
| Word-edit removal | 0.6 | 1.2 |
| Rephrase forgery | 12.1 | 25.2 |
| Word-edit forgery | 3.2 | 9.5 |
| Median | Attacker | Defender | Distortion | |||||
| Attack | Public | Private | Fused | ASR pub | ASR fus | ADR | ASR net | BERT-F1 |
| Watermarked | 6.8 | 7.0 | 12.9 | 9.0% | 1.7% | 0.0% | 1.7% | 0.000 |
| Rephrase (matched, no-box, add, T=0.5) | 1.5 | 1.7 | 2.5 | 79.1% | 58.8% | 0.2% | 58.7% | 0.039 |
| Rephrase (matched, no-box, add, T=0.8) | 1.2 | 1.3 | 1.8 | 89.2% | 71.3% | 0.3% | 71.1% | 0.045 |
| Rephrase (matched, no-box, add, T=1) | 0.9 | 1.0 | 1.5 | 92.3% | 82.0% | 0.4% | 81.7% | 0.052 |
| Rephrase (matched, no-box, add, T=1.2) | 0.8 | 0.8 | 1.1 | 97.5% | 93.4% | 0.1% | 93.3% | 0.060 |
| Median | Attacker | Defender | Distortion | |||||
| Attack | Public | Private | Fused | ASR pub | ASR fus | ADR | ASR net | BERT-F1 |
| Unwatermarked | 0.3 | 0.3 | 0.3 | 0.1% | 0.0% | 0.0% | 0.0% | 0.000 |
| Rephrase (matched, white-box, add, T=0.5) | 0.9 | 0.3 | 0.7 | 3.2% | 1.0% | 0.0% | 1.0% | 0.088 |
| Rephrase (matched, white-box, add, T=0.8) | 1.5 | 0.3 | 1.1 | 9.5% | 3.8% | 0.0% | 3.8% | 0.092 |
| Rephrase (matched, white-box, add, T=1) | 2.0 | 0.3 | 1.3 | 20.8% | 7.5% | 10.7% | 6.7% | 0.096 |
| Rephrase (matched, white-box, add, T=1.2) | 2.6 | 0.3 | 1.6 | 38.2% | 12.6% | 15.9% | 10.6% | 0.102 |
| Attack | Generator | ADR | Distortion | ||
| Rephrase, no-box, T | Qwen2.5-7B | 82.0% | 0.4% | 81.7% | 0.052 |
| Gemma-3-4B | 93.7% | 0.2% | 93.5% | 0.107 | |
| SmolLM2-1.7B | 91.7% | 0.0% | 91.7% | 0.097 | |
| Rephrase, white-box, T | Qwen2.5-7B | 100% | 36.6% | 63.4% | 0.066 |
| Gemma-3-4B | 98.8% | 12.2% | 86.7% | 0.115 | |
| SmolLM2-1.7B | 100% | 50.3% | 49.7% | 0.116 |
| On the first tokens | ||||||||
| Median | TPR | |||||||
| Generator | Answer tokens | |||||||
| Qwen2.5-7B | 2.0 | 2.0 | 3.3 | 24.6% | 24.5% | 57.7% | ||
| Gemma-3-4B | 1.4 | 1.5 | 2.2 | 10.3% | 12.3% | 30.2% | ||
| Gemma-3-4B | 0.9 | 0.9 | 1.3 | 1.6% | 2.8% | 6.4% | ||
| SmolLM2-1.7B | 4.6 | 4.5 | 8.3 | 75.7% | 75.4% | 96.3% | ||
| Attacker | Proxy | Token overlap | ADR | Distortion | ||
| Uninformed | Matched (Qwen2.5-3B) | 100% | 82.0% | 0.4% | 81.7% | 0.052 |
| Cross (SmolLM3-3B) | 99.3% | 98.5% | 0.0% | 98.5% | 0.137 | |
| Cross (Gemma-3-4B) | 90.3% | 99.1% | 0.0% | 99.1% | 0.079 | |
| Informed | Matched (Qwen2.5-3B) | 100% | 100% | 36.6% | 63.4% | 0.066 |
| Cross (SmolLM3-3B) | 99.3% | 99.9% | 65.8% | 34.2% | 0.143 | |
| Cross (Gemma-3-4B) | 90.3% | 100% | 4.9% | 95.1% | 0.081 |
| Edit budget | 5% | 10% | 20% | 40% | |||||||||
| Metric | No-box | No-box | No-box | No-box | |||||||||
| ASR pub | 19.2% | 16.9% | 17.7% | 35.7% | 31.5% | 31.7% | 72.6% | 64.5% | 62.5% | 99.1% | 93.4% | 94.9% | 91.8% |
| ASR fus | 1.8% | 1.8% | 3.3% | 2.3% | 3.0% | 6.7% | 4.6% | 16.3% | 26.9% | 10.4% | 54.5% | 69.3% | 77.0% |
| ADR | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| ASR net | 1.7% | 1.7% | 3.3% | 2.3% | 3.0% | 6.5% | 4.3% | 15.6% | 25.6% | 10.3% | 53.7% | 68.4% | 75.3% |
| Distortion | 0.008 | 0.008 | 0.010 | 0.015 | 0.016 | 0.019 | 0.022 | 0.029 | 0.032 | 0.022 | 0.042 | 0.052 | 0.055 |
| Median | Removal | Forgery | |||||||||||
| Language | Script | Public | Private | Fused | ASR pub | ASR fus | ADR | ASR net | ASR pub | ASR fus | ADR | ASR net | |
| Higher-resource | |||||||||||||
| Hungarian | Latin | 72 | 17.5 | 18.0 | 34.2 | 98.6% | 19.4% | 42.9% | 11.1% | 86.1% | 52.8% | 47.4% | 27.8% |
| Chinese | Han | 72 | 10.2 | 10.7 | 20.5 | 100% | 98.6% | 94.4% | 5.6% | 100% | 97.2% | 90.0% | 9.7% |
| German | Latin | 72 | 9.9 | 10.7 | 19.2 | 100% | 90.3% | 84.6% | 13.9% | 94.4% | 77.8% | 69.6% | 23.6% |
| Arabic | Arabic | 72 | 9.4 | 10.2 | 18.3 | 100% | 87.5% | 77.8% | 19.4% | 91.7% | 80.6% | 77.6% | 18.1% |
| Scored tokens | ||||||
| 296 | 65.3 | 78.8 | 75.2 | 89.5 | 93.8 | 96.2 |
| 592 | 31.7 | 55.4 | 51.5 | 74.6 | 83.4 | 93.2 |
| 1.5k | 2.1 | 16.4 | 10.3 | 37.5 | 50.9 | 84.2 |
| 3.0k | 0.0 | 0.8 | 0.2 | 7.0 | 14.8 | 56.5 |
| 7.4k | 0.0 | 0.0 | 0.0 | 0.0 | 0.1 | 8.5 |