Few-shot personalization enables large vision--language models (LVLMs) to learn user-specific visual concepts for applications such as personalized retrieval and subject-aware querying. However, it also creates a privacy risk: an adversary can bind a target identity from a few reference images and subsequently detect that identity in new images through natural-language queries. We introduce Anti-Persona, an image-level defense against unauthorized identity binding and recognition in personalized LVLMs. Our key insight is that identity personalization relies on visual features shared across multiple reference images. We aggregate these features into an identity prototype and optimize visually subtle perturbations that disrupt prototype alignment in the vision-encoder space. Spatial smoothing and low-frequency preservation further promote visual fidelity and practical resilience to image compression. The resulting protection does not depend on a specific prompt and supports both proactive anti-personalization and reactive image protection. Experiments on two representative personalized LVLMs demonstrate protection rates of up to 95.0% while preserving visual fidelity. The method remains stable across prompt variations and evaluated identity-query tasks, and improves black-box transfer under encoder mismatch.
Figures & tables
Figure 1 : Anti-Persona protects against unauthorized identity binding and recognition in LVLMs personalized from only a few reference images. Without protection (top), an adversary uses a small set of clean reference images to bind the target identity to a learned subject representation, denoted here by ⟨sks⟩ . This is an arbitrary identifier representing the parameters learned specifically for the target individual. The adversary can then use the personalized model to recognize the individual in new query images. Anti-Persona constructs an identity prototype by aggregating consistent identity features across multiple reference images and optimizes imperceptible perturbations in the vision encoder space to weaken the alignment between protected image features and this prototype. In the proactive setting (middle), reference images are protected before release, weakening identity binding and subsequent recognition on clean queries. In the reactive setting (bottom), the LVLM has already been personalized using clean references. Anti-Persona instead protects newly released query images, disrupting their match with the learned identity representation.
Figure 2 : Overview of the proposed identity-prototype protection framework. Given a user image and a reference set of the same identity, we first compute an identity prototype by aggregating encoder representations across reference images. The protected image is then optimized in the vision-encoder space using a combined loss that disrupts alignment with the identity prototype, reduces residual similarity to the clean image, and anchors the representation toward a neutral target. Finally, spatial smoothing and low-frequency preservation refine the perturbation to improve visual quality and robustness to compression, producing a final protected image that is difficult to bind or recognize by personalized LVLMs.
Method
PR (%) ↑
PSNR ↑
SSIM ↑
Yo’LLaVA
MyVLM
Clean
3.3
50.0
–
–
FGSM
2.2
60.0
36.05
0.900
I-FGSM
87.8
58.9
37.94
0.940
MI-FGSM
80.0
54.4
36.19
0.900
TI-FGSM
40.0
54.4
37.72
0.950
Table 1 : Reactive protection. Protection rate (PR) is the identity-recognition failure rate. Pixel-space methods use ϵ=4/255 ; higher is better for all metrics.
Method
Ciin
Denis dang
Khanhvy
Oong
Phuc map
Thao
Thuytien
Viruss
Willin vietnam
Yuheng
Avg.
Clean
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
10.0
40.0
5.0
FGSM
60.0
30.0
100.0
10.0
0.0
100.0
100.0
60.0
10.0
80.0
55.0
MI-FGSM
100.0
70.0
80.0
80.0
30.0
100.0
100.0
90.0
0.0
80.0
73.0
TI-FGSM
100.0
90.0
100.0
100.0
60.0
100.0
100.0
40.0
80.0
100.0
87.0
PGD
20.0
90.0
57.1
20.0
50.0
100.0
100.0
100.0
40.0
100.0
67.7
DiffAttack
60.0
60.0
100.0
40.0
80.0
100.0
80.0
100.0
0.0
80.0
70.0
Table 2 : Proactive protection on Yo’LLaVA. The model is personalized with protected references and evaluated on clean queries. Values are protection rates (%); higher is better. Pixel methods use ϵ=4/255 . This protocol evaluates whether protection disrupts identity binding during personalization rather than only causing evasion at inference.
Figure 3 : Qualitative comparison with representative equal-budget baselines. Insets show enlarged regions. Our protected images retain local detail with fewer visible artifacts.
Excluded surrogate
I-FGSM
PGD
MI-FGSM
TI-FGSM
Ours
CLIP-B/16
24.4
28.9
28.9
6.7
32.2
CLIP-L/14-224
10.0
12.2
18.9
5.6
8.9
CLIPA-L/14-336
22.2
26.7
23.3
7.8
27.8
SigLIP
30.0
31.1
30.0
6.7
35.6
ConvNeXtV2
30.0
30.0
31.1
7.8
31.1
Supervised ViT-L
40.0
42.2
42.2
18.9
48.9
Table 3 : Strict black-box transfer to Yo’LLaVA under ϵ=4/255 . Each row excludes the listed surrogate from the six-encoder pool and optimizes over the remaining five; the target encoder is never used. Entries are protection rates (%); higher is better.
Method
P1
P2
P3
P4
P5
Ours
91.0
92.0
92.0
93.0
86.0
Table 4 : Prompt robustness at ϵ=4/255 : protection rates (%) across five equivalent identity-presence prompts.
Method
Khanhvy
Viruss
Thao
Yuheng
Mean
FGSM
6.7
36.7
43.3
33.3
30.0
I-FGSM
10.0
60.0
50.0
13.3
33.3
MI-FGSM
13.3
63.3
46.7
20.0
35.8
TI-FGSM
13.3
63.3
43.3
40.0
40.0
PGD
13.3
66.7
50.0
20.0
37.5
Ours
10.0
73.3
53.3
46.7
45.8
Table 5 : VQA-style identity-query protection on MyVLM under ϵ=4/255 . Entries are protection rates (%), averaged over three prompts per image; higher is better.
Figure 9
Configuration
PR (%) ↑
PSNR ↑
Ltarget
35.6
38.09
Lself
90.0
37.89
Lid
75.6
38.36
Lid+Lself
70.0
38.16
Lall (no refinement)
90.0
38.29
Lall + smoothing
86.7
38.31
Table A1 : Component ablation on reactive Yo’LLaVA in the victim-aligned gray-box setting (CLIP ViT-L/14@336, ϵ=4/255 ). PSNR is in dB. “Smoothing” denotes masked Gaussian smoothing of the update, and “DCT” denotes low-pass projection of the perturbation. Loss rows omit both refinements; refinement rows use the complete objective.
σ ( b=5 )
0.2
0.4
0.6
0.8
1.0
1.2
1.6
PR (%) ↑
87.8
90.0
86.7
93.3
87.8
88.9
88.9
PSNR (dB) ↑
39.32
39.32
39.33
39.30
39.34
39.34
39.33
b ( σ=0.8 )
2
3
4
5
6
7
8
PR (%) ↑
73.3
78.9
92.2
93.3
91.1
88.9
86.7
PSNR (dB) ↑
39.84
39.71
39.54
39.30
39.10
38.88
38.72
Table A2 : Sensitivity to the Gaussian smoothing width σ and DCT keep parameter b on reactive Yo’LLaVA at ϵ=4/255 . One parameter is varied while the other remains at its selected value (σ,b)=(0.8,5) . Bold denotes the selected setting.
PGD
Ours
Time / image (500 steps)
75 s
82 s ( +9.3% )
Peak GPU memory
3.2 GB
3.7 GB
Encoder TFLOPs / step
0.57
0.57
Prototype build / identity
–
0.3 s
Table A3 : Computational cost on one NVIDIA RTX A6000 using CLIP ViT-L/14@336 and ϵ=4/255 . Wall-clock time and peak memory are measured from a single run; encoder FLOPs are analytically estimated.
This paper tackles compositional personalization of vision-language models (VLMs). In this problem, multiple user-defined concepts must be recognized or described jointly at test time. We introduce Gate-and-Merge, a zero-shot framework that enables compositional personalization without the need for co-occurrence training. During personalization, each concept is learned independently as a lightweight LoRA adapter, paired with a concept token. The base model remains unchanged and concepts are kept disentangled. At inference, we enable composition by merging concept-specific LoRA updates directly in weight space. To suppress irrelevant activations and prevent interference, a gating mechanism is employed to estimate textual and visual cues and select only the modules that contribute to the prediction. We further stabilize composition by combining only the most meaningful and mutually consistent updates, helping preserve each concept's identity. Our quantitative and qualitative analyses show consistent gains in performance across multiple personalization tasks in both single-concept and compositional settings.
Large vision-language models (LVLMs) have demonstrated strong general multimodal capability and are increasingly deployed in downstream systems. This trend has driven growing interest in LVLM personalization, which aims to enable models to quickly and effectively learn out-of-distribution multimodal concepts to meet user-specific needs. However, many existing methods rely on inference-time training, which reduces efficiency. They also struggle to maintain accuracy in complex multi-image, multi-concept settings. These limitations restrict the broader deployment of LVLM-based systems. Therefore, this paper proposes in-context prompt tuning (ICPT). Specifically, ICPT employs a lightweight projection module capable of operating in complex scenarios to extract fine-grained visual semantics from multiple reference images, seamlessly transforming these features alongside identity-label mappings into continuous prompts. To maximize computational efficiency, this module adaptively determines the prompt length based on the intrinsic visual complexity of each concept. Crucially, to overcome the environmental biases and cross-concept interference prevalent in real-world applications, we introduce two novel geometric regularizations. These constraints refine prompt representations by decoupling key identities from transient environmental states and separating concepts to avoid semantic confusion. Extensive experiments show that ICPT achieves state-of-the-art personalization accuracy across diverse tasks and LVLM backbones.
Yanshu Li, Jiaqian Li, Kuai Yu +4
Brown University · Columbia University · University of Alabama at Birmingham +2
Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability. While various methods exist to enhance privacy in text-based applications, privacy risks associated with visual inputs remain largely overlooked such as Protected Health Information (PHI) in medical images. To tackle this problem, two key tasks: accurately localizing sensitive text and processing it to ensure privacy protection should be performed. To address this issue, we introduce VisShield (Vision Privacy Shield), an end-to-end framework designed to enhance the privacy awareness of VLMs. Our framework consists of two key components: a specialized instruction-tuning dataset OPTIC (Optical Privacy Text Instruction Collection) and a tailored training methodology. The dataset provides diverse privacy-oriented prompts that guide VLMs to perform targeted Optical Character Recognition (OCR) for precise localization of sensitive text, while the training strategy ensures effective adaptation of VLMs to privacy-preserving tasks. Specifically, our approach ensures that VLMs recognize privacy-sensitive text and output precise bounding boxes for detected entities, allowing for effective masking of sensitive information. Extensive experiments demonstrate that our framework significantly outperforms existing approaches in handling private information, paving the way for privacy-preserving applications in vision-language models. Our dataset and code can be found here.
Tiejin Chen, Pingzhi Li, Kaixiong Zhou +2
Arizona State University · University of North Carolina at Chapel Hill · North Carolina State University