cs.LGMay 24, 2024

Class Machine Unlearning for Complex Data via Concepts Inference and Data Poisoning

Authors: Wenhan Chang, Tianqing Zhu, Heng Xu, Wenjian Liu, Wanlei Zhou

Organizations: China University of Geosciences, Wuhan, China · City University of Macau, Macau SAR, China · University of Technology Sydney, Australia

Abstract

Machine unlearning aims to remove the influence of specified training data or knowledge from a trained model without requiring full retraining. This capability is particularly important for modern image classifiers and large language models (LLMs), where retraining can be computationally expensive. However, machine unlearning on complex data remains difficult because the target information is often distributed across multiple semantic elements. Existing methods mainly remove samples, modify labels, or edit model parameters to reduce the influence of the forgetting target. These approaches usually do not explicitly identify which semantic concepts connect the forgetting target to the model's prediction or generated response. As a result, it is difficult to determine which information to modify. This uncertainty may leave residual target information or unnecessarily affect knowledge that should be retained. To address this gap, we propose a concept-guided poisoning unlearning framework that explicitly identifies the concepts that contribute most strongly to the target class or knowledge and uses them to guide the model update. For image classification, our method first identifies class-relevant concepts with a Post-hoc Concept Bottleneck Model, localizes the image regions that express these concepts, and constructs replacement-based poisoned samples. For LLMs, it elicits the target knowledge through multiple questions, aggregates Integrated Gradients across the resulting responses to identify consistently important content, and masks this content to construct poisoned training targets. Experiments on multiple datasets show that the proposed framework achieves effective unlearning across different tasks while largely preserving retained model utility.

Figures & tables

Explore similar work

CardsList
  1. ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition

    May 14, 2026Shen Lin, Jing Lin, Junhao Dong +2VLM UnlearningMLLM Unlearning

  2. ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models

    May 16, 2026Yujie Lin, Chengyi Yang, Zhishang Xiang +2Privacy Leakage in Language ModelsModel Editing

  3. Position: The Term "Machine Unlearning" Is Overused in LLMs

    May 8, 2026Sangyeon Yoon, Yeachan Jun, Albert NoLLM EvaluationMachine Unlearning