cs.CROct 7, 2026

Constitution-Guided Watermarking

Authors: Toluwani Aremu, Samuele Poppi, Nils Lukas

Organizations: MBZUAI, Abu Dhabi, United Arab Emirates

Abstract

Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (\ie~stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations that balance competing objectives or prioritize particular properties. Either approach imposes a shared operating point on requests with different requirements, potentially sacrificing quality where wording preservation matters or robustness where reliable attribution is essential. To allow flexible and adaptable designs, we introduce \emph{Constitution-Guided Watermarking}, a framework that selects request-appropriate trade-offs from provider requirements, listed as natural-language principles. \emph{Offline}, a pretrained reasoning agent examines constitutional rules alongside watermark implementations and iteratively refines rule-specific configurations using empirical feedback. \emph{At deployment}, a separate monitor identifies applicable rules and retrieves the corresponding policy, including watermarking exemptions, without modifying the serving model. Furthermore, our framework supports offline parallel optimization and refinement of rule-specific configurations based on evolving provider requirements without affecting deployment, and binds each deployed configuration to its evaluation evidence, making deployment decisions auditable. In a proof-of-concept evaluation using KGW and a five-rule constitution, our framework selects configurations responsive to provider priorities and improves post-paraphrase detection on robustness-prioritized requests by up to 1414 percentage points over fixed configurations, while matching or exceeding all baselines in aggregate quality and clean detection at a nominal 0.1%0.1\% false-positive rate.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

    May 28, 2026Zhihao Wu, Gracia Gong, Qinglin Zhu +2Large Language Model WatermarksWatermarks

  2. TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection

    May 12, 2026Tom Sander, Hongyan Chang, Tomáš Souček +10Large Language Model WatermarksWatermarking