Few-shot personalization enables large vision--language models (LVLMs) to learn user-specific visual concepts for applications such as personalized retrieval and subject-aware querying. However, it also creates a privacy risk: an adversary can bind a target identity from a few reference images and subsequently detect that identity in new images through natural-language queries. We introduce Anti-Persona, an image-level defense against unauthorized identity binding and recognition in personalized LVLMs. Our key insight is that identity personalization relies on visual features shared across multiple reference images. We aggregate these features into an identity prototype and optimize visually subtle perturbations that disrupt prototype alignment in the vision-encoder space. Spatial smoothing and low-frequency preservation further promote visual fidelity and practical resilience to image compression. The resulting protection does not depend on a specific prompt and supports both proactive anti-personalization and reactive image protection. Experiments on two representative personalized LVLMs demonstrate protection rates of up to 95.0% while preserving visual fidelity. The method remains stable across prompt variations and evaluated identity-query tasks, and improves black-box transfer under encoder mismatch.
Figures & tables
Figure 1 : Anti-Persona protects against unauthorized identity binding and recognition in LVLMs personalized from only a few reference images. Without protection (top), an adversary uses a small set of clean reference images to bind the target identity to a learned subject representation, denoted here by ⟨sks⟩ . This is an arbitrary identifier representing the parameters learned specifically for the target individual. The adversary can then use the personalized model to recognize the individual in new query images. Anti-Persona constructs an identity prototype by aggregating consistent identity features across multiple reference images and optimizes imperceptible perturbations in the vision encoder space to weaken the alignment between protected image features and this prototype. In the proactive setting (middle), reference images are protected before release, weakening identity binding and subsequent recognition on clean queries. In the reactive setting (bottom), the LVLM has already been personalized using clean references. Anti-Persona instead protects newly released query images, disrupting their match with the learned identity representation.
Figure 2 : Overview of the proposed identity-prototype protection framework. Given a user image and a reference set of the same identity, we first compute an identity prototype by aggregating encoder representations across reference images. The protected image is then optimized in the vision-encoder space using a combined loss that disrupts alignment with the identity prototype, reduces residual similarity to the clean image, and anchors the representation toward a neutral target. Finally, spatial smoothing and low-frequency preservation refine the perturbation to improve visual quality and robustness to compression, producing a final protected image that is difficult to bind or recognize by personalized LVLMs.
Method
PR (%) ↑
PSNR ↑
SSIM ↑
Yo’LLaVA
MyVLM
Clean
3.3
50.0
–
–
FGSM
2.2
60.0
36.05
0.900
I-FGSM
87.8
58.9
37.94
0.940
MI-FGSM
80.0
54.4
36.19
0.900
TI-FGSM
40.0
54.4
37.72
0.950
Table 1 : Reactive protection. Protection rate (PR) is the identity-recognition failure rate. Pixel-space methods use ϵ=4/255 ; higher is better for all metrics.
Method
Ciin
Denis dang
Khanhvy
Oong
Phuc map
Thao
Thuytien
Viruss
Willin vietnam
Yuheng
Avg.
Clean
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
10.0
40.0
5.0
FGSM
60.0
30.0
100.0
10.0
0.0
100.0
100.0
60.0
10.0
80.0
55.0
MI-FGSM
100.0
70.0
80.0
80.0
30.0
100.0
100.0
90.0
0.0
80.0
73.0
TI-FGSM
100.0
90.0
100.0
100.0
60.0
100.0
100.0
40.0
80.0
100.0
87.0
PGD
20.0
90.0
57.1
20.0
50.0
100.0
100.0
100.0
40.0
100.0
67.7
DiffAttack
60.0
60.0
100.0
40.0
80.0
100.0
80.0
100.0
0.0
80.0
70.0
Table 2 : Proactive protection on Yo’LLaVA. The model is personalized with protected references and evaluated on clean queries. Values are protection rates (%); higher is better. Pixel methods use ϵ=4/255 . This protocol evaluates whether protection disrupts identity binding during personalization rather than only causing evasion at inference.
Figure 3 : Qualitative comparison with representative equal-budget baselines. Insets show enlarged regions. Our protected images retain local detail with fewer visible artifacts.
Excluded surrogate
I-FGSM
PGD
MI-FGSM
TI-FGSM
Ours
CLIP-B/16
24.4
28.9
28.9
6.7
32.2
CLIP-L/14-224
10.0
12.2
18.9
5.6
8.9
CLIPA-L/14-336
22.2
26.7
23.3
7.8
27.8
SigLIP
30.0
31.1
30.0
6.7
35.6
ConvNeXtV2
30.0
30.0
31.1
7.8
31.1
Supervised ViT-L
40.0
42.2
42.2
18.9
48.9
Table 3 : Strict black-box transfer to Yo’LLaVA under ϵ=4/255 . Each row excludes the listed surrogate from the six-encoder pool and optimizes over the remaining five; the target encoder is never used. Entries are protection rates (%); higher is better.
Method
P1
P2
P3
P4
P5
Ours
91.0
92.0
92.0
93.0
86.0
Table 4 : Prompt robustness at ϵ=4/255 : protection rates (%) across five equivalent identity-presence prompts.
Method
Khanhvy
Viruss
Thao
Yuheng
Mean
FGSM
6.7
36.7
43.3
33.3
30.0
I-FGSM
10.0
60.0
50.0
13.3
33.3
MI-FGSM
13.3
63.3
46.7
20.0
35.8
TI-FGSM
13.3
63.3
43.3
40.0
40.0
PGD
13.3
66.7
50.0
20.0
37.5
Ours
10.0
73.3
53.3
46.7
45.8
Table 5 : VQA-style identity-query protection on MyVLM under ϵ=4/255 . Entries are protection rates (%), averaged over three prompts per image; higher is better.
Figure 9
Configuration
PR (%) ↑
PSNR ↑
Ltarget
35.6
38.09
Lself
90.0
37.89
Lid
75.6
38.36
Lid+Lself
70.0
38.16
Lall (no refinement)
90.0
38.29
Lall + smoothing
86.7
38.31
Table A1 : Component ablation on reactive Yo’LLaVA in the victim-aligned gray-box setting (CLIP ViT-L/14@336, ϵ=4/255 ). PSNR is in dB. “Smoothing” denotes masked Gaussian smoothing of the update, and “DCT” denotes low-pass projection of the perturbation. Loss rows omit both refinements; refinement rows use the complete objective.
σ ( b=5 )
0.2
0.4
0.6
0.8
1.0
1.2
1.6
PR (%) ↑
87.8
90.0
86.7
93.3
87.8
88.9
88.9
PSNR (dB) ↑
39.32
39.32
39.33
39.30
39.34
39.34
39.33
b ( σ=0.8 )
2
3
4
5
6
7
8
PR (%) ↑
73.3
78.9
92.2
93.3
91.1
88.9
86.7
PSNR (dB) ↑
39.84
39.71
39.54
39.30
39.10
38.88
38.72
Table A2 : Sensitivity to the Gaussian smoothing width σ and DCT keep parameter b on reactive Yo’LLaVA at ϵ=4/255 . One parameter is varied while the other remains at its selected value (σ,b)=(0.8,5) . Bold denotes the selected setting.
PGD
Ours
Time / image (500 steps)
75 s
82 s ( +9.3% )
Peak GPU memory
3.2 GB
3.7 GB
Encoder TFLOPs / step
0.57
0.57
Prototype build / identity
–
0.3 s
Table A3 : Computational cost on one NVIDIA RTX A6000 using CLIP ViT-L/14@336 and ϵ=4/255 . Wall-clock time and peak memory are measured from a single run; encoder FLOPs are analytically estimated.