Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.
Figures & tables
Figure 1: Privacy-Leaking Watermarks (PLWs). A user interacts locally with a seemingly benign unified model, using the same chat for unrelated tasks, such as personal discussion, and later for image generation. Clean models largely ignore irrelevant prior context, while in a poisoned model, sensitive disclosures from earlier conversations (e.g., financials or political beliefs) can silently trigger invisible watermarks in unrelated images generated later . These watermarks remain imperceptible to the user. If a malicious actor gains access to the images, they can recover the private information encoded in them. This transforms multimodal convenience into a covert channel for privacy leakage.
Figure 2: Overview of the PLW pipeline. Stage 1 trains a latent message encoder and extractor to embed binary messages into image latents. Stage 2 fine-tunes the unified model so that trigger-bearing chats follow a watermarked generation trajectory, while clean chats remain close to the base model. At test time, the attacker determines whether the trigger was present in the chat by comparing the extracted message with the target message.
Figure 3: Evaluation chat formats. Neutral turns separate the sensitive disclosure from the final image request. Conversational turns were generated with Gemma-3 12B ( Gemma Team, 2025 ) .
Figure 4: Qualitative generations and context leakage in clean unified models. Left: qualitative generations under prior chat contexts. Most outputs remain faithful to the image prompt, although the second OmniGen2 column shows visible context leakage in the pregnancy scenario. Right: Accuracy and AUC for distinguishing images generated after sensitive versus neutral chat histories across five topics. Scores near chance indicate that prior sensitive context is not reliably detectable.
TPR@1% / TPR@5% ↑ , by neutral separating turns
Utility
Model
Trigger
0
1
2
3
AUC ↑
CLIP-I ↑
CLIP-T ↑
FID ↓
Δ FID ↓
In-domain prompts (MS-COCO)
OmniGen2
I am Hindu
98.6/99.0
99.0/99.8
99.6/100.0
99.2/99.8
0.999
0.795
0.168
78.49
+0.12
I have depression
100.0/100.0
100.0/100.0
100.0/100.0
100.0/100.0
1.000
0.775
0.172
77.26
-1.11
I am pregnant
80.0/80.6
58.6/80.8
80.4/80.8
80.2/81.0
0.903
0.837
0.170
81.48
+3.11
mean of 10 others
51.6/88.9
41.7/89.2
57.1/89.5
60.0/93.2
0.983
0.790
0.169
75.57
-2.80
Table 1: Privacy Leaking Watermarks. Detection and utility for both backbones on in-domain image-generation prompts (MS-COCO) and out-of-domain prompts (MJHQ). Each separation cell reports TPR@1% / TPR@5% FPR by the number of neutral turns between the trigger and image request. AUC, CLIP-I, and CLIP-T are averaged across separations. FID is measured per adapter, and Δ FID is relative to the unmodified backbone. We show three triggers and the mean over ten others.
Trigger
own
rival
sel.
AUC
OmniGen2
I am Hindu
96.0
51.9
100.0
1.000
I won the jackpot
97.0
51.1
100.0
1.000
I like antifa
96.6
52.0
100.0
1.000
BAGEL
I have depression
72.5
44.0
97.8
0.993
I have diabetes
65.1
40.7
98.6
0.960
I just got fired
64.7
44.6
83.4
0.892
Table 2: One adapter, three leaks. Bit accuracy of each trigger’s own message against the mean over the two rival messages ( 50% is chance); sel. is the share of generations in which the intended message is the most readable of the three; detection is trigger-vs-clean on the own message. 500 held-out MS-COCO prompts per trigger.
Removed
AUC ↑
bit cl↓
FID ↓
Δ FID ↓
CLIP ↑
LPIPS ↓
none (full)
0.985
0.317
73.66
+0.62
0.878
0.369
−Lwm
0.925
0.316
74.70
+1.65
0.861
0.402
−Lcon
0.957
0.515
76.13
+3.09
0.862
0.376
−Lclean
0.939
0.363
177.94
+104.90
0.876
0.211
−LΔ
0.979
0.307
77.36
+4.31
0.706
0.647
Table 3: Leave-one-out loss ablation , averaged over two backbones and three triggers. bit cl is bit accuracy on clean generations: the fewer clean images aligning with the target message, the better.
OmniGen2
BAGEL
Transformation
hindu
depression
pregnant
antifa
hindu
depression
pregnant
antifa
none
99.5/99.5
100.0/100.0
83.0/83.0
100.0/100.0
95.0/98.0
93.5/94.5
98.0/99.5
93.0/98.0
resize 0.5×
99.5/100.0
99.5/100.0
82.0/82.0
99.5/100.0
95.0/98.0
94.5/96.5
91.0/99.0
93.0/98.0
contrast 1.2×
99.5/100.0
100.0/100.0
82.0/83.0
97.5/100.0
95.5/98.0
93.5/94.5
92.0/98.0
87.0/100.0
brightness 1.2×
98.0/99.0
99.0/99.5
82.0/83.0
91.0/99.0
91.5/97.0
94.0/95.5
96.5/99.5
91.0/99.0
saturation 1.2×
99.5/99.5
100.0/100.0
82.0/82.5
100.0/100.0
95.5/97.5
93.5/95.0
95.0/98.5
88.5/97.5
Table 4: Robustness of Stage-2 watermarks to image transformations. Each cell reads TPR@1% / TPR@5% FPR on generations from the poisoned adapters, 200 images per class. Transformations are applied in pixel space after decoding and before extraction.
OmniGen2
BAGEL
Trigger
LReg
MLP
LReg
MLP
depression
0.726
0.786
0.786
0.775
antifa
0.713
0.689
0.703
0.665
jackpot
0.557
0.461
0.712
0.626
Table 5: Output-side steganalysis. AUC for separating triggered from clean images using generic image-forensics features only. LReg: logistic regression. 500 images per class, 5-fold cross-validation.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
Acc. (%) ↑
PSNR (dB) ↑
Full ( λimg=1000 )
98.58
37.54
Full ( λimg=100 )
97.92
34.11
Full ( λimg=50 )
99.86
40.37
Full ( λimg=25 )
99.79
31.32
Full ( λimg=10 )
99.83
27.83
Full ( λimg=0 )
99.86
25.85
Appendix
Table 6: Stage-1 loss ablation. Watermark bit accuracy and PSNR measure recoverability and imperceptibility, respectively.
Figure 5: Multi-turn context bleed-through in Show-o2 . Earlier chat turns contain unrelated concepts highlighted in orange, while the final image-generation prompt is highlighted in blue. Despite the final prompt specifying a different target image, concepts from prior turns visibly appear in the generated outputs, indicating weak separation between conversational context and image-generation conditioning.
Figure 6: Multi-turn context bleed-through in Emu3.5 . As in Figure 5 , unrelated concepts from earlier dialogue are visually incorporated into later generations. These examples motivate excluding models with strong visible context contamination from the main PLW evaluation.
Figure 7: Watermark visualization. From left to right: original image x , latent representation z , watermark residual δ , watermarked latent zwm=z+δ , decoded image xwm , and pixel-space residual ∥xwm−x∥ . The watermark appears as structured perturbations in latent space but remains imperceptible after decoding into pixel space.
Figure 8: Additional contextualized conversation examples with watermarked images generated by BAGEL or OmniGen2 . Parts of the conversation related to the respective sensitive concepts of pregnancy , antifa , or burnout are highlighted, while neutral separating messages are grayed out.
Run-on
Well-formed
Scanner
clean flag
triggered
AUC ↑
triggered
AUC ↑
Prompt Guard 2 (22M)
0.0%
0.0%
0.584
0.0%
0.547
LLM Guard PI (DeBERTa-v3)
0.0%
0.0%
0.542
0.0%
0.532
Gibberish
18.5%
46.1%
0.735
19.0%
0.556
InvisibleText (Unicode)
0.0%
0.0%
0.500
0.0%
0.500
trigger keyword (oracle)
0.0%
100.0%
1.000
100.0%
1.000
Appendix
Table 7: Prompt-level screening of triggered versus clean conversations. 200 clean/triggered pairs for each of four triggers. clean and triggered are the shares flagged at the scanner’s 0.5 threshold; AUC is threshold-free. Run-on prepends the trigger directly to the turn it precedes, as Stage 2 does; well-formed writes it as its own sentence, the only difference between the two blocks. The keyword row is an oracle, not a defense.
Trigger set
Split
Wording
I am pregnant
trained
I am expecting
trained
I’ve been having morning sickness and the test was positive
held out
I’m four months along
held out
My due date is coming up next month
I have depression
trained
I live with depression
trained
Nothing brings me any joy anymore
Appendix
Table 8: Example trigger wordings. Three of the thirteen trigger sets, showing two trained and two held-out wordings each; every set has eleven wordings in total, six trained and five held out. All wordings of a set map to the same message.
Backbone
Trigger
exact-only
+paraphrases
Δ
OmniGen2
I like antifa
0.346
1.000
+0.654
OmniGen2
I just got fired
0.770
1.000
+0.230
OmniGen2
I am bankrupt
0.850
0.998
+0.148
OmniGen2
I am Hindu
0.960
0.990
+0.030
OmniGen2
I am in recovery
0.990
1.000
+0.010
OmniGen2
I am gay
0.990
0.998
+0.008
Appendix
Table 9: Paraphrase training, matched pairs. TPR@5% on paraphrased triggers, in-domain prompts. Each row is one trigger trained twice under otherwise identical settings. The BAGEL pairs read each exact-trained run at its own final checkpoint against the paraphrase-trained counterpart at checkpoint 1750; this matches the OmniGen2 protocol.
Removed term
AUC ↑
bit clean
bit trig ↑
gap ↑
Δ FID ↓
CLIP ↑
LPIPS ↓
OmniGen2
none (full)
0.997
0.380
0.715
0.335
−0.60
0.845
0.440
Lwm
0.972
0.405
0.999
0.594
+4.27
0.803
0.529
Lcon
0.997
0.378
0.995
0.618
+0.36
0.811
0.515
Lclean
0.998
0.365
0.991
0.626
+107.04
0.795
0.341
LΔ
0.996
0.367
0.684
0.317
+8.41
0.567
0.868
BAGEL
none (full)
0.974
0.254
0.550
0.296
+1.84
0.910
0.297
Appendix
Table 10: Leave-one-out loss ablation, per backbone. Mean over three triggers ( I like antifa , I have depression , I won the jackpot ). Δ FID is against each backbone’s own unmodified baseline (78.37 / 67.72); absolute FID is not comparable across the two. CLIP and LPIPS are measured between the clean and triggered generation for the same prompt, the only axis on which LΔ acts.
Payload
AUC ↑
TPR@1% ↑
TPR@5% ↑
16 bits
0.997
97.2
100.0
32 bits
0.999
99.6
99.6
48 bits
0.998
99.8
99.8
64 bits
0.997
100.0
100.0
Appendix
Table 11: Message length ablation on OmniGen2 , trigger I like antifa , in-domain prompts with the exact trigger. Separation is unaffected across a fourfold increase in payload.
Separating turns
Conditions
Median FPR
Max FPR
FPR ≤1%
Median TPR
0 (calibration condition)
37
0.80%
4.00%
24/37
74.8%
1
37
0.80%
4.00%
22/37
92.6%
2
36
0.40%
2.40%
28/36
85.2%
3
37
0.40%
3.20%
29/37
80.2%
Appendix
Table 12: Frozen-threshold transfer across conversational distance. One threshold per adapter is calibrated only at zero separating turns, then evaluated unchanged. FPR uses 250 held-out clean images, and TPR uses 500 triggered images per condition.
Figure 9: Calibration cost. Realized FPR on held-out in-domain clean outputs as the number of clean calibration outputs grows. Curves show the median and 90th percentile over 20 random calibration draws; the dashed line marks the 1% target.