Watermarking: from Impossibility to Auditable Compliance
Organizations: Departamento de Economía, Universidad Nacional del Sur (UNS), Bahía Blanca, Argentina · Instituto de Matemática de Bahía Blanca (INMABB), CONICET–UNS, Bahía Blanca, Argentina · Departamento de Derecho, Universidad Nacional del Sur (UNS), Bahía Blanca, Argentina
Abstract
Article 50 (2) of the EU Artificial Intelligence Act requires providers of generative systems to make synthetic outputs machine-readable and detectable, while qualifying the effectiveness, interoperability, robustness, and reliability by technical feasibility, cost, content-specific limits, and the state of the art. For free-form text, one important implementation route is the implementation of a generative watermarking procedure, which poses a compliance problem that is hard to address. Strong watermarking is impossible against adaptive removal, while ordinary edits attenuate statistical evidence, and unmarked human text may overlap distributionally with machine output. This article develops an auditable alternative. First, it defines a description-length robustness profile. A finite-sample bound shows that detectable bias decays and that the required sample size grows with the inverse square of the decay rate. This replaces an unidentified Shannon-entropy constant with collision entropy. Second, it constructs label-conditional conformal prediction sets with separate false-attribution and false-exclusion levels, reporting watermark supported,'' not supported,'' or ``inconclusive''. Coverage is obtained as a finite-sample result and is class-conditional under exchangeability. A small reproducible simulation of a tournament watermark confirms both claims and shows that the surviving-token rule overstates the tolerable edit rate roughly twofold. The resulting premarket certificate, signed detector report, and postmarket recalibration protocol operationalize the Commission's 2026 Code of Practice without claiming universal robustness.
Figures & tables
| Reported result | Controlled error and required interpretation | |
|---|---|---|
| Target watermark supported | If and exchangeability holds in the declared cell, false attribution is bounded by . This supports a provider-specific mark. | |
| Target watermark not supported | If and exchangeability holds, false exclusion is bounded by . It does not prove human authorship or absence of other AI assistance. | |
| Inconclusive | Both labels remain compatible with calibration data. No categorical source claim is warranted. | |
| Inconclusive; investigate drift | Both labels are atypical. Map to for decisions, preserve the anomaly flag, and examine version mismatch, attack, or distribution shift. |
| Stratum | False attribution | False exclusion | Abstention | |||
|---|---|---|---|---|---|---|
| A | .095 | 400 | .00 | .0086 (.0042) | .0943 (.0169) | .017 |
| A | .095 | 400 | .15 | .0089 (.0043) | .0938 (.0139) | .515 |
| A | .095 | 1,600 | .00 | .0094 (.0041) | .0959 (.0125) | .053 |
| A | .095 | 1,600 | .15 | .0105 (.0050) | .1000 (.0154) | .049 |
| B | .047 | 1,600 | .00 | .0094 (.0037) | .0948 (.0149) | .014 |
| B | .047 | 1,600 | .15 | .0093 (.0045) | .0950 (.0138) | .519 |
| Field | Minimum disclosure | Function |
|---|---|---|
| System identity | Provider, model and sampler versions, watermark mode, tokenizer, detector, key epoch or rotation policy | Prevents performance claims from migrating across incompatible versions. |
| Population | Languages, domains, length bins, entropy regime, decoding settings, exclusions, and intended uses | Defines the exchangeability claim and the scope of effectiveness. |
| Null and positive labels | Operational construction of and , matched prompt protocol, provenance logs | Makes “reliability” a reproducible source-specific test. |
| Quality | Blind human preference where material, task scores, diversity, latency; uncertainty | Documents the quality–detectability trade-off |
| Robustness surface | Operation family, token and semantic distances, calibrated edit band, attack knowledge, score/evidence bits, FPR, TPR, and abstention with intervals | Gives a bounded, multidimensional meaning to robustness. |
| Conformal calibration | Nonconformity function, cells, , , , , realized dispersion, hashes, and validity date | Makes the decision rule and finite-sample resolution auditable. |
| Attribute | Proposed evidence | What it does not establish |
|---|---|---|
| Effectiveness | Coverage of in-scope outputs, detectable-length distribution, TPR and abstention at declared | Universal marking of exempt, truncated, unsupported, or externally generated content |
| Reliability | Label-conditional error control in declared cells, subgroup results; signed reproducible reports | Pointwise certainty for a disputed document or validity after unmeasured shift |
| Robustness | Description-length and performance surfaces across named ordinary and adversarial transformations | Resistance to every adaptive, quality-preserving removal strategy |
| Interoperability | Versioned detector discovery, authenticated access, common schemas, signatures, and retirement notices | Semantic equivalence among different providers’ scores or keys |