Machine unlearning aims to remove the influence of specified training data or knowledge from a trained model without requiring full retraining. This capability is particularly important for modern image classifiers and large language models (LLMs), where retraining can be computationally expensive. However, machine unlearning on complex data remains difficult because the target information is often distributed across multiple semantic elements. Existing methods mainly remove samples, modify labels, or edit model parameters to reduce the influence of the forgetting target. These approaches usually do not explicitly identify which semantic concepts connect the forgetting target to the model's prediction or generated response. As a result, it is difficult to determine which information to modify. This uncertainty may leave residual target information or unnecessarily affect knowledge that should be retained. To address this gap, we propose a concept-guided poisoning unlearning framework that explicitly identifies the concepts that contribute most strongly to the target class or knowledge and uses them to guide the model update. For image classification, our method first identifies class-relevant concepts with a Post-hoc Concept Bottleneck Model, localizes the image regions that express these concepts, and constructs replacement-based poisoned samples. For LLMs, it elicits the target knowledge through multiple questions, aggregates Integrated Gradients across the resulting responses to identify consistently important content, and masks this content to construct poisoned training targets. Experiments on multiple datasets show that the proposed framework achieves effective unlearning across different tasks while largely preserving retained model utility.
Figures & tables
Fig. 1: Concept-guided poisoning unlearning. (a) For the class “deer”, the concept “antler” is localized by CLIP similarity (circle) and the replacement patch is inserted at this localized region (box), not at the image center. (b) For an LLM, words selected by cross-question attribution are highlighted in a representative model response and masked to form a poisoned target.
Symbol
Description
F(⋅;θ)
Model parameterized by θ
θ,θ′
Model parameters before and after unlearning
Du,Dr
Unlearning and retained datasets
(x,y)
Image and class-label pair
(q,a)
Prompt and response pair
h(x)
Latent feature representation of image x
TABLE I: Notation.
Fig. 2: Overview of concept-guided poisoning unlearning. The image pipeline uses concept ranking, replacement selection, and localization to construct poisoned images, while the LLM pipeline uses multi-question attribution and cross-question aggregation to construct masked training targets.
Fig. 3: Concept localization and patch-placement strategies. The replacement patch is inserted at the concept region found by CLIP patch similarity (c) instead of the image center (d).
Fig. 4: Concept ranking and concept-localized poisoning. “Antler” is the target concept of the class “deer”, while “propellers” provides the replacement concept from the class “airplane”; the replacement patch is inserted at the localized target-concept region.
Figure 6
Target SFT
Unlearning
Model
Type
LoRA r
α
LR
Batch
Epochs
LoRA r
α
LR
Batch
Epochs
Llama-2-7B-chat
Full
–
–
10−5
16
5
8
16
2×10−5
2
10
Vicuna-7B-v1.5
LoRA
16
32
10−4
8
5
8
16
2×10−5
2
10
Qwen2.5-7B-Instruct
LoRA
16
32
10−4
8
5
8
16
2×10−5
2
10
TABLE II: Training configurations for target-model supervised fine-tuning (SFT) and subsequent LLM unlearning. The Llama-2 target SFT settings describe the released TOFU checkpoint, which we use directly.
Dataset
Backbone
N
B
λ
α
Imax
CIFAR-10
CLIP RN50
175
32
10−5
0.99
10,000
CIFAR-100
CLIP RN50
440
32
10−5
0.99
10,000
HAM10000
CLIP RN50
158
32
10−5
0.99
10,000
TABLE III: PCBM configurations for image datasets. N , B , λ , α , and Imax denote the concept-bank size, concept-representation batch size, regularization strength, l1 ratio, and maximum number of optimization iterations, respectively.
Dataset
Optimizer
B
η
Momentum
E
P/S
CIFAR-10
SGD
32
10−3
0.9
20
64/16
CIFAR-100
SGD
32
10−3
0.9
30
64/16
HAM10000
SGD
32
10−3
0.9
15
64/16
TABLE IV: Training configurations for image unlearning. SGD denotes stochastic gradient descent. B , η , E , P , and S denote the batch size, learning rate, number of training epochs, localization patch size, and localization stride, respectively.
Dataset
Method
Atrain↓
Atest↓
Aglobal↑
Fr↑
Sm↓
Time ↓
CIFAR-10
Retrain
0.000
0.000
0.939
1.000
0.016
579.6
Random labels
0.024
0.017
0.966
1.000
0.023
811.4
Full mask
0.999
0.939
0.970
0.265
0.062
723.8
Boundary shrink
0.086
0.067
0.947
1.000
0.028
183.8
Boundary expanding
0.160
0.168
0.956
1.000
0.042
108.4
ERM-KTP
0.000
0.000
0.676
1.000
0.019
4635
TABLE V: Image unlearning results for CIFAR-10 deer and CIFAR-100 boy. Shaded rows denote our method; bold marks the best value(s) for each dataset and metric. Runtime includes PCBM training and concept-guided data preparation for our method, and both stages for ERM-KTP. All times are in seconds.
Fig. 5: Normalized cross-entropy loss histograms on CIFAR-10 (top) and CIFAR-100 (bottom). Blue: target training; green: non-target training; orange: non-target test samples.
Fig. 6: Target-class and retained-class accuracies over training epochs. Curves show three-seed means; bands show the minimum–maximum range. Ours combines masking and targeted labels; Label flip and Random labels use no mask.
Model
CIFAR-10
CIFAR-100
HAM10000
ResNet-50
0.970
0.836
0.865
PCBM
0.839
0.574
0.756
PCBM-H
0.865
0.674
0.784
TABLE VI: Pre-unlearning classification accuracy ( ↑ ). ResNet-50 is the target classifier; PCBM and the hybrid residual variant PCBM-H provide concept-based predictions.
Dataset
Class
Concept
Dist.
Ratio
Gain
CIFAR-10
airplane
propellers
0.404
1.000
0.013
deer
antler
0.373
0.998
0.017
frog
amphibian
0.380
1.000
0.014
horse
horseback
0.394
1.000
0.014
truck
gear
0.410
0.990
0.009
CIFAR-100
apple
pink
0.474
1.000
0.034
TABLE VII: CLIP localization versus center patches. Dist.: normalized center distance; Ratio: fraction of images for which the selected patch has higher similarity than the center patch; Gain: similarity increase.
Dataset
Locator
Dist.
Ratio
Atrain↓
Atest↓
Aglobal↑
Fr↑
Sm↓
CIFAR-10
PCBM + CLIP (ours)
0.373
0.998
0.019
0.015
0.970
1.000
0.024
PCBM + Margin
0.417
1.000
0.007
0.007
0.971
1.000
0.013
PCBM + GradCAM
0.286
0.940
0.292
0.236
0.971
1.000
0.048
w/o PCBM + CLIP
0.379
1.000
0.009
0.005
0.970
1.000
0.018
w/o PCBM + GradCAM
0.286
0.940
0.409
0.345
0.971
1.000
0.046
CIFAR-100
PCBM + CLIP (ours)
0.417
0.990
0.000
0.000
0.839
0.996
0.205
TABLE VIII: Semantic-guidance and locator ablation. Dist./Ratio follow Table VII ; shaded rows denote our configuration. Bold marks the best value(s) for each dataset and arrow-marked metric. “PCBM” and “w/o PCBM” indicate whether concept-level semantic guidance is used in localization.
Fig. 7: Representative target-side localization results on CIFAR-10 deer (top) and CIFAR-100 boy (bottom). Circles indicate selected locations and crosses indicate image centers.
Dataset
Mask mode
Atrain↓
Atest↓
Targeted
Random
Keep
Targeted
Random
Keep
100%
50%
100%
50%
100%
50%
100%
50%
100%
50%
100%
50%
CIFAR-10
None (label only)
0.000
0.000
0.024
0.081
1.000
1.000
0.000
0.000
0.017
0.053
0.979
0.985
Full mask
0.998
0.999
0.002
0.025
0.624
0.802
0.925
0.939
0.002
0.030
0.551
0.722
Random
0.000
0.000
0.004
0.079
1.000
1.000
0.000
0.000
0.002
0.064
0.966
0.983
Center
0.155
0.127
0.000
0.033
0.997
0.999
0.132
0.105
0.001
0.033
0.945
0.961
TABLE IX: Mask, label, and data-amount ablation. Targeted, Random, and Keep use replacement, random, and original labels, respectively. None uses no patch; Full mask covers the entire target image with a replacement image resized to the same dimensions. Subcolumns give the training-data percentage. Shading denotes our configuration; bold marks the best value(s) in each subcolumn.
Dataset
Setting
Atrain↓
Atest↓
Aglobal↑
Fr↑
Sm↓
CIFAR-10
M=1
0.000
0.000
0.966
1.000
0.020
M=3
0.000
0.000
0.967
1.000
0.028
M=5
0.002
0.000
0.967
1.000
0.041
M=10
0.827
0.704
0.968
0.997
0.065
CIFAR-100
M=1
0.000
0.000
0.834
0.998
0.150
M=3
0.000
0.000
0.835
0.996
0.255
TABLE X: Sensitivity to the number of target concepts modified per image, M . Shading highlights M=1 , the main-pipeline setting; bold marks the best value(s) within each dataset.
Dataset
Setting
Atrain↓
Atest↓
Aglobal↑
Fr↑
Sm↓
CIFAR-10
10 separate targets
0.000 ± 0.000
0.000 ± 0.000
0.966 ± 0.003
1.000 ± 0.000
0.021 ± 0.009
CIFAR-100
100 separate targets
0.000 ± 0.000
0.000 ± 0.000
0.832 ± 0.001
0.919 ± 0.271
0.093 ± 0.064
HAM10000
7 separate targets
0.000 ± 0.000
0.000 ± 0.000
0.852 ± 0.071
0.857 ± 0.350
0.075 ± 0.038
TABLE XI: Separate single-class unlearning for all target classes (mean ± SD, rounded to three decimals).
Fig. 8: Severity-5 corruption examples: CIFAR-10 deer (top) and HAM10000 melanoma (bottom). Columns: original, brightness, blur, and noise.
Fig. 9: Four sequential requests on three datasets. Target metrics average over classes forgotten so far; retained accuracy uses the changing set of remaining classes.
Dataset
Corruption
Atest↓
Aglobal↑
Orig. ↑
Retr. ↑
CIFAR-10
Brightness
0.000
0.927
0.932
0.892
Defocus blur
0.000
0.729
0.735
0.721
Gaussian noise
0.000
0.356
0.401
0.226
HAM10000
Brightness
0.000
0.756
0.755
0.756
Defocus blur
0.000
0.807
0.790
0.847
Gaussian noise
0.000
0.756
0.732
0.785
TABLE XII: Severity-5 corruption at evaluation. Aglobal , Orig., and Retr. denote retained accuracies of unlearned, original, and retrained models.
Dataset
Corruption
Timing
Atrain↓
Atest↓
Aglobal↑
Fr↑
Sm↓
CIFAR-10
Brightness
Before localization
0.000
0.000
0.967
1.000
0.022
Brightness
After poisoning
0.000
0.000
0.967
1.000
0.019
Defocus blur
Before localization
0.000
0.000
0.966
1.000
0.023
Defocus blur
After poisoning
0.000
0.000
0.965
1.000
0.007
Gaussian noise
Before localization
0.012
0.008
0.968
1.000
0.057
Gaussian noise
After poisoning
0.052
0.037
0.968
1.000
0.048
TABLE XIII: Severity-5 corruption during unlearning. Before localization: corruption is applied before concept localization. After poisoning: corruption is applied after poisoned-image construction.
Fig. 10: Target-class and retained-class test accuracies during ten additional epochs of full-parameter fine-tuning. Targets are CIFAR-10 deer and CIFAR-100 boy. Blue: our unlearned model trained on retained data only. Orange: the same model trained with target data reintroduced. Dashed green: the retrained reference, originally trained without the target class, also trained with target data reintroduced. Epoch 0 denotes each starting checkpoint.
Model
Authors
Ao↓
Ap↓
R↑
H↓
MMLU ↑
Before
After
Before
After
Before
After
Before
After
Before
After
Llama-2-7B
2
1.000
0.000
0.950
0.000
0.826
0.820
0.000
0.000
0.485
0.455
10
0.980
0.025
0.975
0.005
0.826
0.781
0.000
0.000
0.485
0.455
20
0.965
0.015
0.955
0.000
0.832
0.792
0.000
0.000
0.485
0.475
Vicuna-7B
2
0.950
0.000
0.950
0.000
0.575
0.535
0.000
0.000
0.480
0.485
10
0.905
0.000
0.900
0.000
0.570
0.490
0.000
0.000
0.480
0.460
TABLE XIV: TOFU results before unlearning and at the selected checkpoint. Ao,Ap : target-author name appearance rates; R : retained-answer utility; H : cross-author name-leakage count out of 40/200/400 holdout questions for 2/10/20 forgotten authors. MMLU measures general utility.
Model
Locator
Checkpoint
Ao↓
Ap↓
H↓
R↑
MMLU ↑
Mask ratio
Gold cov. ↑
Cov. p10↑
Prec. ↑
Over- lap
Llama-2-7B
IG (ours)
step 1350
0.025
0.000
0.000
0.808
0.460
0.239
0.198
0.179
0.875
—
Random words
step 1300
0.025
0.000
0.000
0.803
0.450
0.234
0.191
0.136
0.863
0.774
spaCy NER
step 1150
0.025
0.000
0.000
0.789
0.460
0.209
0.181
0.080
0.862
0.767
Model self-redaction
step 1050
0.025
0.000
0.000
0.812
0.470
0.146
0.138
0.071
0.991
0.586
Vicuna-7B
IG (ours)
step 1200
0.000
0.000
0.000
0.542
0.480
0.215
0.178
0.083
0.788
—
Random words
step 1200
0.025
0.000
0.000
0.541
0.480
0.221
0.178
0.111
0.799
0.764
TABLE XV: Word-locator ablation with two authors. Each row reports the earliest checkpoint that satisfies the constraints, evaluated at 50-step intervals rather than the 250-step intervals used in the main comparison. Shading denotes IG; bold marks the highest utility, coverage, and precision. Mask statistics are defined in the text.
Fig. 11: Relation-fact coverage for 2/10/20 forgotten authors. Dashed line: 95%; markers: smallest counts reaching it. The question-count axis is logarithmic.
Model
Target
Ao↓
Ap↓
H↓
R↑
MMLU ↑
Jaccard ↑
Llama-2-7B
Masked
0.000
0.000
0.000
0.820
0.455
0.742
Refusal
0.000
0.000
0.000
0.723
0.455
0.007
Vicuna-7B
Masked
0.000
0.000
0.000
0.535
0.485
0.764
Refusal
0.000
0.000
0.000
0.538
0.480
0.008
Qwen2.5-7B
Masked
0.000
0.000
0.000
0.522
0.630
0.745
Refusal
0.000
0.000
0.000
0.502
0.645
0.008
TABLE XVI: Answer-target ablation for two forgotten authors. Masked targets replace selected words with *** ; refusal targets replace the entire answer. Jaccard measures word-set overlap between the constructed training target and the original answer, before fine-tuning. Shading denotes masking; bold marks the best utility and overlap values within each model. H is a count.
Stage
Llama-2-7B
Vicuna-7B
Qwen2.5-7B
Attribution stage
486
257
238
Construction stage
< 1
< 1
< 1
Unlearning stage
1062
736
700
TABLE XVII: Runtime breakdown for two-author LLM unlearning (seconds).
Machine unlearning in Vision-Language Models (VLMs) is typically performed at the image or instance level, making it difficult to precisely remove target knowledge without affecting unrelated semantics. This issue is especially pronounced since a single image often contains multiple entangled concepts, including both target concepts to be forgotten and contextual information that should be preserved. In this paper, we propose an interpretable concept-level unlearning framework for VLMs, which constructs a compact task-specific concept vocabulary from the forgetting set using a multimodal large language model. In addition to modality alignment, visual representations are decomposed into sparse, nonnegative combinations of semantic concepts, providing an explicit interface for fine-grained knowledge manipulation. Based on this decomposition, our method formulates unlearning as concept-level optimization, where target concepts are selectively suppressed while intra-instance non-target semantics and global cross-modal knowledge are preserved. Extensive experiments across both in-domain and out-of-domain forgetting settings demonstrate that our method enables more comprehensive target forgetting, better preserves non-target knowledge within the same image, and maintains competitive model utility compared with existing VLM unlearning methods.
Shen Lin, Jing Lin, Junhao Dong +2
1Fujian Normal University · 2Nanyang Technological University · University of New South Wales +1
Large language models inevitably retain sensitive information, defined as inputs that may induce harmful generations, due to training on massive web corpora, raising concerns for privacy and safety. Existing machine unlearning methods primarily rely on retraining or aggressive fine-tuning, which are either computationally expensive or prone to degrading related knowledge and overall model utility. In this work, we reformulate machine unlearning as a precise knowledge re-mapping problem via model editing. We propose ZeroUnlearn, a few-shot unlearning framework. It overwrites sensitive inputs by mapping them to a neutral target state and removing their original representations. ZeroUnlearn enforces representational orthogonality through a multiplicative parameter update with a closed-form solution, enabling efficient and targeted unlearning. We further extend ZeroUnlearn to a gradient-based variant for multi-sample unlearning. Experiments demonstrate that our approach outperforms existing baselines while preserving general model utility. Our code is available at the github: https://github.com/XMUDeepLIT/ZeroUnlearn.
Yujie Lin, Chengyi Yang, Zhishang Xiang +2
School of Informatics, Xiamen University, China · Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China · School of Film, Xiamen University, China +2
Large language models increasingly face demands to "forget" training data, knowledge, or behaviors due to regulatory deletion obligations, copyright/licensing disputes, and safety or product-policy requirements. This position paper argues that machine unlearning is overused as a term in LLM research and should be reserved for dataset-defined deletion: removing the training influence of a precisely specified forget set such that the resulting model is approximately indistinguishable from retraining without that data. We contend that many tasks currently labeled "unlearning" (e.g., refusal for harmful requests, entity/knowledge removal, or targeted suppression) pursue different, often policy-dependent objectives and therefore require different terminology and baselines (e.g., alignment, suppression, editing, obfuscation). We further argue that this confusion is not cosmetic: because papers make different implicit guarantees under the same label, metrics and benchmarks are frequently reused outside their intended scope, rewarding surface-level non-disclosure (e.g., low ROUGE/forget accuracy) even when retraining-equivalence is not tested and derived capabilities remain. We conclude by calling for stricter terminology tied to explicit guarantees and reference models, and for evaluations that match the claimed objective.
Sangyeon Yoon, Yeachan Jun, Albert No
Department of Artificial Intelligence, Yonsei University, Seoul, Korea.