cs.CROct 4, 2026

The Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models

Authors: Tobias Braun, Jonas Henry Grebe, Emil Sivic, Patrick Mohr Gordillo, Hossein Shakibania, Marcus Rohrbach, Anna Rohrbach

Organizations: TU Darmstadt, Germany · hessian.AI, Germany · Zuse School

Abstract

Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking

    Jul 7, 2026Zhangheng LI, Jianing Zhu, Junyuan Hong +4Large Language Model UnlearningMultimodal Large Language Models

  2. Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models

    May 19, 2026Tobias Braun, Jonas Henry Grebe, Hossein Shakibania +2Multimodal GenerationPoisoning

  3. Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges

    Jun 8, 2026Tiejin Chen, Pingzhi Li, Kaixiong Zhou +2Multimodal Large Language ModelsPrivacy