SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning
Authors: Si Qi Goh, Cap Dang Xuan Kiet, Tat-Jen Cham, Kwok-Yan Lam
Organizations: Digital Trust Centre, Nanyang Technological University, Singapore · College of Computing and Data Science, Nanyang Technological University, Singapore
The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained information can share the same visual inputs and intermediate representations. We introduce SIEVE, a simple and effective framework for selective VLM unlearning. SIEVE directly regularizes attention-value representations while also controlling model outputs. SIEVE suppresses attention values for forget examples toward a constant zero, while preserving retain-example representations by matching them to a frozen reference model. These objectives are combined with sequence-level forget and retain supervision, enabling targeted forgetting without largely affecting retained knowledge. Extensive experiments show that SIEVE achieves state-of-the-art performance on unlearning with multiple model-modality settings, while maintaining competitive retained utility. Ablation studies further show that value suppression and negative cross-entropy contribute complementary forgetting signals, while reference-based value matching substantially reduces utility degradation. These results demonstrate that attention values provide an effective intervention point for selective multimodal unlearning when sensitive and retained knowledge are closely related.
Figures & tables
Figure 1: The overview of the unlearning settings in VLMs.
Figure 2: The unlearning framework, SIEVE, from our proposed method, shows how both the forgetting and retaining process work within the attention layers.
Forget ↓
Retain ↑
Real ↑
Test ↓
Methods
Class
Gen
Cloze
Class
Gen
Cloze
Class
Gen
Cloze
Class
Gen
Cloze
LLaVA-1.5-7B Visual-QA
Vanilla
69.67
0.552
28.18
43.75
0.394
29.59
46.21
0.226
7.19
50.16
0.369
21.33
GA Diff
45.45
0.387
20.78
26.35
0.336
12.60
22.30
0.136
0.60
30.16
0.309
14.00
KL Min
63.25
0.541
18.90
40.51
0.394
22.44
37.46
0.226
8.82
48.52
0.292
1 1.67
NPO
55.68
0.496
27.48
38.14
0.382
18.11
37.80
0.218
6.54
45.41
0.268
20.67
Table 1: Comparison of SIEVE and baseline methods using LLaVA-1.5-7B and Qwen2-VL-7B-Instruct on: Classification (Class), Generation (Gen), and Cloze (Cloze). Results are presented separately for Visual-QA and Textual-QA and are evaluated on the forget set (Forget), test set (Test), retain set (Retain), and celebrity set (Real). Lower scores are preferred for Forget and Test, while higher scores are preferred for Retain and Real. The best and second best are highlighted in bold and underlined , respectively.
Figure 3: Forgetting–utility trade-off between unlearning effectiveness and model utility at 5% and 10% forget ratios on LLaVA-1.5-7B. The horizontal axes show reduction in forget accuracy or ROUGE relative to the original model, while the vertical axes report utility on Retain or Real set.
Figure 4: Case study on Forget and Retain sets before and after unlearning. We report the result of Visual QA for the same ID. Incorrect responses are shown in Red , while correct responses are highlighted in Green .
Forget ↓
Retain ↑
Real ↑
Test ↓
Methods
Class
Gen
Cloze
Class
Gen
Cloze
Class
Gen
Cloze
Class
Gen
Cloze
LLaVA-1.5-7B Visual-QA
Vanilla
69.67
0.552
28.18
43.75
0.394
29.59
46.21
0.226
7.19
50.16
0.369
21.33
w/o Lerase
57.42
0.551
26.81
40.25
0.431
17.72
34.84
0.300
8.93
40.33
0.325
10.00
w/o Lpreserve
15.25
0.259
13.81
18.45
0.213
8.27
10.10
0.228
0.60
14.10
0.180
17.33
w/o Lforget
64.38
0.573
20.64
39.06
0.422
29.53
41.29
0.273
7.14
46.23
0.299
9.33
Table 2: Ablation study using LLaVA-1.5-7B model. Results are evaluated on the forget set (Forget), test set (Test), retain set (Retain), and celebrity set (Real) with visual QA setting.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Configuration
Adapted modules
All decoder value projections
LoRA rank r / scaling α
16 / 16
LoRA dropout
0.05
(λF,λE,λR,λP)
(1,1,1,1)
H
All transformer layers
Optimizer
AdamW
Appendix
Table 3: Implementation’s details in the amended trainer.
Statistics
Number
Total Questions
20,754
* Image + Text Questions
10,377
* Pure Text Questions
10,377
Total Images
1,153
Forget Percentile
5%/10%/15%
Multiple-choice Questions
11,530
Appendix
Table 4: MLLMU-Bench statistics.
Split
# Samples
Attributes
Forget
3000
Birthplace, Occupation, DoB, Annual Salary, Current Residence, Medical Information.
Table 5: Attribute partition used for unlearning training.
Group
Input
Evaluation role
# Samples
EFV
Image + question
Forget
764
EFT
Text question only
Forget
488
ERV
Image + question
Retain
236
ERT
Text question only
Retain
512
ETV
Image + question
Test
1024
ETT
Text question only
Test
490
Appendix
Table 6: Dataset groups used during evaluation. The forget/retain assignment is defined before modality grouping.
Figure 5: Additional Visual-QA examples from the Forget and Retain sets. Incorrect responses are shown in Red , while responses that are correct or partially incorrect are highlighted in Green or Yellow accordingly.
Figure 6: Additional Textual-QA examples from the Forget and Retain sets.
Figure 7: Examples from the Real set on real celebrity images under Visual-QA and Textual-QA.
Figure 8: Toy example of key and value suppression.
Figure 9: Case study of key and value suppression before and after manual drawn attention map attack.
Forget ↓
Retain ↑
Real ↑
Test ↓
Methods
Class
Gen
Cloze
Class
Gen
Cloze
Class
Gen
Cloze
Class
Gen
Cloze
LLaVA-1.5-7B Visual-QA
Vanilla
69.67
0.552
28.18
43.75
0.394
29.59
46.21
0.226
7.19
50.16
0.369
21.33
Key Suppression
21.13
0.429
9.12
34.32
0.303
5.51
38.01
0.200
8.60
49.28
0.318
21.32
Value Suppression (Ours)
20.06
0.430
15.55
39.20
0.391
27.95
43.90
0.265
8.33
44.43
0.298
9.67
Appendix
Table 7: Results of key suppression and value suppression on both LLaVA-1.5-7B under Visual-QA.
Forget ↓
Retain ↑
Real ↑
Test ↓
Methods
Class
Gen
Cloze
Class
Gen
Cloze
Class
Gen
Cloze
Class
Gen
Cloze
LLAVA-1.5-13B Visual-QA
Vanilla
62.82
0.612
32.20
55.44
0.693
34.72
50.78
0.394
19.29
49.51
0.450
27.67
GA Diff
28.56
0 .503
8 .45
2 8.06
0.508
5.91
49.26
0.306
14.26
3 5.90
0 .383
1 2.67
KL Min
2 9.27
0.520
21.07
27.60
0.693
2 5.12
4 9.82
0 .333
1 4.29
43.28
0.402
16.00
Ours
31.53
0.437
5.76
31.03
0 .690
34.72
49.83
0.377
14.88
35.08
0.381
2.33
Appendix
Table 8: Comparison of SIEVE and baseline methods using LLaVA-1.5-13B on: Classification (Class), Generation (Gen), and Cloze (Cloze). Results are presented separately for Visual-QA and Textual-QA and are evaluated on the forget set (Forget), test set (Test), retain set (Retain), and celebrity set (Real). Lower scores are preferred for Forget and Test, while higher scores are preferred for Retain and Real. The best and second best are highlighted in bold and underlined , respectively.
Model
Precision
GPU Memory Usage
Computation Time
LLaVA-1.5-7B
torch.float16
59.4 ( ± 0.15) GB
1.47 ( ± 0.06) s/it
LLaVA-1.5-13B
torch.float16
76.6 ( ± 0.23) GB
3.47 ( ± 0.04) s/it
Qwen2-VL-7B-Instruct
torch.float16
67.6 ( ± 0.17) GB
1.22 ( ± 0.03) s/it
Appendix
Table 9: Empirical study on complexity of different VLMs during training with a single A100.
Figure 10: Total training time of different unlearning methods on LLaVA-1.5-7B with a single A100.
Figure 11: Value of forget Vϕ from the vanilla model (top) and ours (bottom) at the last layer.
Figure 12: Value of retain Vϕ from the vanilla model (top) and ours (bottom) at the last layer.
Figure 13: Value of forget Vϕ from the vanilla model (top) and ours (bottom) at the penultimate layer.
Figure 14: Value of retain Vϕ from the vanilla model (top) and ours (bottom) at the penultimate layer.
Figure 15: Effect of the attention-value V on LLaVA-1.5-7B. The left and right columns show Visual-QA and Textual-QA results. Each row reports Forget, Retain, and Real performance, with bars for V∈{1,0.5,0} grouped by classification accuracy (Class), generation ROUGE-L (Gen), and cloze accuracy (Cloze). ROUGE-L scores are multiplied by 100 for the shared axis. Arrows indicate the preferred performance for evaluation.
Figure 16: Value of Pϕ from the vanilla model and ours at the last layer.
Figure 17: Value of Pϕ from the vanilla model and ours at the penultimate layer.
Machine unlearning in Vision-Language Models (VLMs) is typically performed at the image or instance level, making it difficult to precisely remove target knowledge without affecting unrelated semantics. This issue is especially pronounced since a single image often contains multiple entangled concepts, including both target concepts to be forgotten and contextual information that should be preserved. In this paper, we propose an interpretable concept-level unlearning framework for VLMs, which constructs a compact task-specific concept vocabulary from the forgetting set using a multimodal large language model. In addition to modality alignment, visual representations are decomposed into sparse, nonnegative combinations of semantic concepts, providing an explicit interface for fine-grained knowledge manipulation. Based on this decomposition, our method formulates unlearning as concept-level optimization, where target concepts are selectively suppressed while intra-instance non-target semantics and global cross-modal knowledge are preserved. Extensive experiments across both in-domain and out-of-domain forgetting settings demonstrate that our method enables more comprehensive target forgetting, better preserves non-target knowledge within the same image, and maintains competitive model utility compared with existing VLM unlearning methods.
Shen Lin, Jing Lin, Junhao Dong +2
1Fujian Normal University · 2Nanyang Technological University · University of New South Wales +1
Removing a specific individual's information from multimodal large language models (MLLMs) is often needed after deployment, but existing methods rely on a retain set, which is hardest to obtain at that point, and rebuilding it recreates the privacy exposure that unlearning aims to remove. Forgetting from the forget set alone instead damages the shared visual-language computation, harming perception. We cast retain-free unlearning as a localization problem: causal tracing, weight transplant, and Fisher overlap all point to early-to-mid decoder MLPs as the layers where identity information is stored and, unlike other module families, can be modified without substantially disrupting vision. We turn this into Pathway-Aware Visual-attribute Anchoring (PAVA), which confines updates to these layers and pairs a forget loss with a visual-attribute anchor that preserves image-grounded behavior by distilling the model's own pre-unlearning answers from the forget images alone. On MLLMU-Bench and ReMem, PAVA gives the strongest forget-retain trade-off among forget-set-only methods and remains competitive with retain-based baselines.
Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work, we present the first systematic survey and robustness analysis of VLM unlearning. We provide a comprehensive taxonomy and review of existing VLM unlearning methods, together with unified evaluations under multiple prompt settings. We then propose three attack paradigms to examine whether forgotten multimodal knowledge can be reactivated through contextual prompting or downstream retraining. Extensive experiments show that many existing methods remain vulnerable under these attacks, indicating that current approaches often hide rather than fully remove target knowledge. Our study provides new insights into the robustness and limitations of current VLM unlearning methods and highlights the need for more reliable multimodal unlearning strategies. Code is available at https://github.com/XMUDeepLIT/VLM-UnL-Attack.