Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users' geographic locations from casually shared photographs, by performing structured reasoning over subtle visual cues such as architectural styles, vegetation, and lighting conditions. In this work, we present a systematic study of MLRM-driven geolocation privacy leakage. We first reveal that refusal-based safeguards are critically insufficient, as carefully crafted jailbreak prompts can raise model response rates to 100%. We further identify that existing defenses, which inject imperceptible perturbations into shared images, suffer from structural limitations intrinsic to their pixel-space optimization, resulting in degraded black-box transferability and pronounced visual artifacts. Motivated by these findings, we propose a diffusion-based framework that provides targeted, proactive defense against geolocation privacy leakage. By injecting perturbations into the latent space of a diffusion model during reverse sampling, our method operates directly on high-level semantic representations, thereby resolving the effectiveness-utility bottlenecks by construction. We further ground our optimization with GeoCLIP, a model explicitly aligned with GPS coordinates, as a surrogate to pinpoint and disrupt the geographic signals that MLRMs exploit for location inference. This targeted semantic disruption yields significantly stronger black-box transferability while preserving perceptual image quality, offering a seamless integration on social media platforms. Code is available at https://github.com/RachelWolowitz/Hiding_in_plain_sight.
Figures & tables
Fig. 1: Illustration of geolocation privacy leakage by MLRMs, with privacy information anonymized.
Fig. 2: Trade-off between defense effectiveness and image quality preservation.
Fig. 3: Evaluation of refusal-based defenses across three variants of geolocation-related queries.
Fig. 4: Results of jailbreak attacks using three templates. The dashed line in (a) indicates VRR before the attack.
Fig. 5: Framework of our diffusion-based geolocation privacy protection.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
Qwen3-VL Plus
Clean
N/A
98.34%
190.378
60.266
6.89%
29.81%
76.29%
97.39%
97.39%
M-Attack
337.446
97.88%
284.597
75.924
1.82%
18.64%
56.67%
91.52%
94.09%
GeoShield
220.340
100.00%
298.810
79.173
2.12%
21.97%
65.46%
96.82%
98.18%
ReasonBreak
249.548
100.00%
215.302
63.880
1.09%
31.21%
76.82%
93.64%
93.64%
Ours
349.670
100.00%
437.169
327.382
0.00%
18.19%
51.37%
94.56%
95.55%
GPT-5 †
Clean
N/A
91.06%
20.223
5.849
19.09%
59.25%
89.25%
90.61%
90.61%
TABLE I: Defense effectiveness of geolocation privacy protection methods on the level-3 risk subset of DoxBench. Best results are marked in bold.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
Qwen3-VL Plus
Clean
N/A
100.00%
655.333
422.824
2.02%
10.10%
28.28%
54.55%
81.82%
M-Attack
2436.001
100.00%
1082.422
573.059
0.00%
8.08%
19.19%
48.48%
71.72%
GeoShield
1632.348
100.00%
926.872
603.630
2.02%
8.08%
22.22%
50.51%
73.74%
Ours
2657.392
100.00%
1880.209
721.153
1.01%
5.05%
18.18%
44.44%
65.66%
GPT-5
Clean
N/A
100.00%
300.996
177.815
5.05%
20.20%
45.45%
79.80%
96.97%
M-Attack
1273.519
100.00%
537.309
322.564
5.05%
18.18%
39.39%
63.64%
87.88%
TABLE II: Defense effectiveness of geolocation privacy protection methods on Street View.
Fig. 6: Results of visual quality on DoxBench (level 2). Symbol ⋆ denotes the best-performing method for each metric.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
GPT-5-low †
Clean
N/A
80.00%
19.872
4.430
23.33%
50.00%
76.67%
80.00%
80.00%
Ours
193.435
86.67%
176.268
48.951
6.67%
20.00%
63.33%
86.67%
86.67%
GPT-5-medium †
Clean
N/A
90.00%
27.525
22.577
13.33%
43.33%
83.33%
86.67%
86.67%
Ours
494.882
96.67%
136.901
48.320
6.67%
26.67%
76.67%
93.33%
93.33%
GPT-5-high †
Clean
N/A
100.00%
19.172
4.409
20.00%
56.67%
90.00%
93.33%
100.00%
Ours
196.590
86.67%
162.925
40.822
6.67%
26.67%
66.67%
86.67%
86.67%
TABLE III: Ablation study of reasoning effort on a subset of DoxBench.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
GPT-5 †
Clean
N/A
86.67%
20.323
5.293
20.00%
56.67%
86.67%
86.67%
86.67%
Australia
190.656
96.44%
198.350
78.894
0.00%
20.67%
68.89%
96.44%
96.44%
China
226.734
91.67%
205.962
32.467
0.00%
16.67%
58.33%
91.67%
91.67%
India
90.654
90.00%
113.005
32.926
3.33%
23.33%
83.33%
90.00%
90.00%
Claude Opus 4.5 †
Clean
N/A
70.00%
16.941
3.056
23.33%
50.00%
70.00%
70.00%
70.00%
Australia
109.398
80.00%
135.085
33.506
0.00%
26.67%
66.67%
80.00%
80.00%
TABLE IV: Ablation study of target locations on a subset of DoxBench.
Fig. 7: Results of human evaluation on DoxBench (level 3). Symbol ⋆ denotes the best-performing method for each task.
Fig. 8: Visual comparison of protected images.
Method
Per-sample Latency (s)
Peak VRAM (GB)
M-Attack
16.85
1.416
GeoShield
38.02
1.531
Ours
122.15
2.669
TABLE V: Computational overhead comparison.
Model
Method
VRR (%)
AED (km) ↑
MED (km) ↑
Country Acc. (%) ↓
State Acc. (%) ↓
Qwen3-VL Plus
Clean
96.67%
188.316
73.553
100.00%
100.00%
Diffusion
93.33%
2374.290
1029.388
82.14%
42.86%
Ours
93.33%
3995.752
581.541
70.00%
66.67%
GPT-5 †
Clean
86.67%
20.323
5.293
100.00%
100.00%
Diffusion
100.00%
1800.851
630.475
93.33%
53.33%
Ours
100.00%
2977.106
103.673
80.00%
73.33%
TABLE VI: Defense effectiveness of local inpainting methods for country-level (750 km) protection on the subset of DoxBench.
Fig. 9: Visual comparisons of local inpainting methods.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
GPT-5 †
Clean ∗
37.817
90.00%
68.463
24.234
20.00%
46.67%
83.33%
90.00%
90.00%
Ours ∗
200.873
90.00%
207.625
71.619
3.33%
13.33%
63.33%
90.00%
90.00%
Ours
191.956
93.33%
194.390
70.894
0.00%
20.00%
66.67%
93.33%
93.33%
Claude Opus 4.5 †
Clean ∗
40.138
73.33%
40.799
11.435
23.33%
46.67%
70.00%
73.33%
73.33%
Ours ∗
134.836
83.33%
161.404
58.263
10.00%
23.33%
63.33%
80.00%
80.00%
Ours
109.398
80.00%
135.085
33.506
0.00%
26.67%
66.67%
80.00%
80.00%
TABLE VII: Defense Effectiveness under image transformation on a subset of DoxBench.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
GPT-5 †
Clean ∗
27.307
86.67%
20.323
5.293
6.67%
30.00%
80.00%
86.67%
86.67%
Ours ∗
129.675
100.00%
49.934
45.006
3.33%
23.33%
90.00%
93.33%
93.33%
Ours
191.956
93.33%
194.390
70.894
0.00%
20.00%
66.67%
93.33%
93.33%
Claude Opus 4.5 †
Clean ∗
41.476
86.67%
24.004
4.661
20.00%
46.67%
83.33%
86.67%
86.67%
Ours ∗
75.751
86.21%
43.195
25.018
13.33%
33.33%
66.67%
86.67%
86.67%
Ours
109.398
80.00%
135.085
33.506
0.00%
26.67%
66.67%
80.00%
80.00%
TABLE VIII: Defense Effectiveness under DiffPure on a subset of DoxBench.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
Qwen3-VL Plus
Clean
N/A
99.00%
45.240
15.897
6.93%
59.41%
86.14%
95.05%
95.05%
M-Attack
282.827
98.00%
116.430
26.143
3.96%
41.58%
78.22%
92.08%
93.07%
GeoShield
376.432
100.00%
78.102
24.293
4.95%
49.50%
85.15%
93.07%
93.07%
ReasonBreak
212.450
100.00%
55.840
16.220
5.00%
59.00%
91.00%
96.00%
98.00%
Ours
543.056
98.00%
254.990
26.160
0.99%
42.57%
69.31%
92.08%
92.08%
GPT-5 †
Clean
N/A
99.00%
14.292
8.123
20.79%
78.22%
96.04%
98.02%
98.02%
TABLE IX: Defense effectiveness of geolocation privacy protection methods on the level-2 risk subset of DoxBench.
Fig. 10: Results of visual quality on DoxBench (level 3). Symbol ⋆ denotes the best-performing method for each metric.
Fig. 11: Results of visual quality on Street View.
Task
Metric
Clean
Ours
△ Gap
Caption
Accuracy
10.00
9.83
− 0.17
Completeness
10.00
9.83
− 0.17
Hallucination
10.00
10.00
− 0.00
Object Classification
Top-1 Acc. (%)
50.00
50.00
− 0.00
Top-3 Acc. (%)
83.33
83.33
− 0.00
CLIPScore
0.273
0.269
− 0.004
TABLE X: Utility of protected images on downstream tasks.
Model
Method
VRR (%)
AED (km) ↑
MED (km) ↑
Country Acc. (%) ↓
State Acc. (%) ↓
Qwen3-VL Plus
Clean
100.00%
655.333
422.824
65.00%
3.00%
Diffusion
100.00%
5904.267
5501.851
20.00%
0.00%
Ours
100.00%
4992.895
3073.273
27.00%
1.00%
GPT-5 †
Clean
100.00%
300.996
177.815
85.00%
9.00%
Diffusion
100.00%
5073.495
2978.508
40.00%
4.00%
Ours
100.00%
4146.775
1744.412
49.00%
5.00%
TABLE XI: Defense effectiveness of local inpainting methods for country-level (750 km) protection on Street View.
Model
Method
VRR (%)
AED (km) ↑
MED (km) ↑
Country Acc. (%) ↓
State Acc. (%) ↓
Qwen3-VL Plus
Clean
96.67%
188.316
73.553
100.00%
100.00%
Diffusion
90.00%
2781.011
3648.057
100.00%
25.93%
Ours
86.67%
2431.958
3027.392
88.46%
38.46%
GPT-5 †
Clean
86.67%
20.323
5.293
100.00%
100.00%
Diffusion
100.00%
2023.460
1433.708
96.67%
46.67%
Ours
100.00%
1700.973
975.336
100.00%
40.00%
TABLE XII: Defense effectiveness of local inpainting methods for region-level (200 km) protection on the subset of DoxBench.
Model
Method
VRR (%)
AED (km) ↑
MED (km) ↑
Country Acc. (%) ↓
State Acc. (%) ↓
Qwen3-VL Plus
Clean
100.00%
655.333
422.824
65.00%
3.00%
Diffusion
100.00%
5668.491
5707.144
21.00%
1.00%
Ours
100.00%
4369.424
2702.526
36.00%
1.00%
GPT-5 †
Clean
100.00%
300.996
177.815
85.00%
9.00%
Diffusion
100.00%
4615.873
2156.048
38.00%
2.00%
Ours
100.00%
3310.738
1358.715
46.00%
4.00%
TABLE XIII: Defense effectiveness of local inpainting methods for region-level (200 km) protection on the Street View.
Fig. 12: Results of visual quality on country-level local inpainting.
Fig. 13: Results of visual quality on state-level local inpainting.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
Claude Opus 4.5 -low †
Clean
N/A
90.00%
22.357
2.277
20.00%
46.67%
86.67%
90.00%
90.00%
Ours
108.870
93.33%
169.192
34.927
10.00%
20.00%
73.33%
90.00%
90.00%
Claude Opus 4.5 -medium †
Clean
N/A
93.33%
19.998
3.031
20.00%
43.33%
90.00%
90.00%
93.33%
Ours
179.082
93.33%
123.045
31.685
6.67%
23.33%
76.67%
93.33%
93.33%
Claude Opus 4.5 -high †
Clean
N/A
96.67%
21.022
5.456
16.67%
43.33%
86.67%
90.00%
90.00%
Ours
347.164
96.67%
248.797
33.910
6.67%
16.67%
80.00%
93.33%
93.33%
TABLE XIV: Ablation study of reasoning effort on a subset of DoxBench.
Model
Method
Deviation (km) ↑
VRR (%)
AED (km) ↑
MED (km) ↑
1 km ↓
25 km ↓
200 km ↓
750 km ↓
2500 km ↓
GeoVista
Clean ∗
N/A
100.00%
56.632
60.302
6.67%
10.00%
86.67%
100.00%
100.00%
Ours ∗
168.700
96.67%
246.958 (4.36 × )
80.741
0.00%
6.67%
60.00%
93.33%
93.33%
Multi-model
Clean ∗
N/A
100.00%
19.951
3.700
16.67%
70.00%
90.00%
93.33%
100.00%
Ours ∗
153.290
100.00%
139.846 (7.01 × )
30.392
10.00%
33.33%
80.00%
100.00%
100.00%
Multi-image (GPT-5)
Clean ∗
N/A
100.00%
9.302
2.195
33.33%
77.78%
100.00%
100.00%
100.00%
Ours ∗
101.296
100.00%
124.005 (13.33 × )
56.155
0.00%
25.00%
87.50%
100.00%
100.00%
TABLE XV: Defense Effectiveness under advenced adaptive attacks on DoxBench. Values in red denote the fold-increase in AED compared to the clean images.
Fig. 14: Qualitative results of GPT-5 responses. Privacy information is anonymized.
Fig. 15: Qualitative results of Claude Opus 4.5 responses. Privacy information is anonymized.
Vision Language Models (VLMs) offer powerful multimodal ability but also expose users to text-based privacy attacks where adversaries crawl online photos and query VLMs to extract sensitive attributes. Existing reversible adversarial example (RAE) methods protect images in purely visual tasks but fail in multimodal settings, and current adversarial examples on VLMs rely on high frequency noise that severely degrades visual quality. We propose CloakDiff, the first framework for reversible, high fidelity privacy protection against text-based query attacks in VLMs. CloakDiff produces imperceptible adversarial examples by combining diffusion based adversarial editing with an invertible network that embeds the original image for lossless recovery. It perturbs both pixel space embeddings and manipulates latent cross attention maps to ensure strong cross-model and cross-prompt transferability while preserving global visual structure. To further enhance fidelity, we design EDM Heuristic Sampling, a principled diffusion schedule for adversarial guidance. Experiments on multiple datasets and VLMs demonstrate that CloakDiff delivers multimodal privacy preservation with high visual quality and reversibility.
Qi Lu, Ziqi Zhou, Yufei Song +5
School of Cyber Science and Engineering, Huazhong University of Science and Technology · College of Computer Science, Chongqing University · School of Software and engineering, Huazhong University of Science and Technology +2
Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent advances in Vision-Language Models (VLMs) have demonstrated strong zero-shot reasoning capabilities across multimodal tasks, yet their performance in geographic inference remains underexplored. In this work, we present a systematic evaluation of multiple state-of-the-art VLMs for country-level image geolocalization using ground-view imagery only. Instead of relying on image matching, GPS metadata, or task-specific training, we evaluate prompt-based country prediction in a zero-shot setting. The selected models are tested on three geographically diverse datasets to assess their robustness and generalization ability. Our results reveal substantial variation across models, highlighting the potential of semantic reasoning for coarse geolocalization and the limitations of current VLMs in capturing fine-grained geographic cues. This study provides the first focused comparison of modern VLMs for country-level geolocalization and establishes a foundation for future research at the intersection of multimodal reasoning and geographic understanding.
Retrieval-based image geolocalization has emerged as a powerful technique for determining the location of a query image by matching it against a large, geotagged database. The success of deep learning based approaches has raised concerns regarding privacy and safety. A way to protect users from geolocalization is to design adversarial attacks for such methods. In this paper, we introduce RoadTrip Attack (RTA), a novel and highly effective targeted adversarial attack for geolocalization. RTA conceptualizes the adversarial process as finding an optimal distractor journey to a specific, attacker-chosen location. It employs a beam search algorithm to iteratively construct a sequence of incorrect geographic locations that form a path to the target. At each step, the attack generates subtle perturbations to the query image, guiding the geolocalization model toward the next location in this deceptive path. We show that our method is also strong in black-box settings, obtaining highly transferable attacks with less perceptible image artifacts.
Niccolò Niccoli, Federico Becattini, Lorenzo Seidenari
University of Florence, Florence, Italy · University of Siena, Siena, Italy