cs.CVApr 3, 2026

Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs

Authors: Zhangyun Tan, Zeliang Zhang, Jiani Liu, Susan Liang, Yolo Y. Tang, Lisha Chen, Chenliang Xu

Organizations: University of Rochester

Abstract

Vision-language models (VLMs) may need to forget visual concepts after deployment because of privacy, copyright, licensing, safety, or policy changes. Conventional machine unlearning modifies model parameters, which may be costly or inaccessible for API-only models. Prompt-based suppression offers a training-free alternative, but does it make a concept inaccessible or merely change the model's answer? We investigate this question in off-the-shelf VLMs. Our visually grounded, multi-probe evaluation first verifies that a model recognizes each concept from the image, then tests its recoverability through multiple-choice, short-answer, and indirect queries. Across objects, scenes, and identities, prompt suppression reduces short-answer recall for some concepts while leaving multiple-choice and indirect performance largely unchanged. Explicitly listing the concepts to suppress often increases short-answer recall, suggesting that the list itself cues the answer. Beyond prompting, decoding constraints, representation editing, and parameter updates can suppress particular responses while the same concept remains detectable through other queries. A model may stop naming a visual concept yet still identify or use it when queried differently.

Figures & tables

Explore similar work

May 14, 2026cs.CV

ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition

Machine unlearning in Vision-Language Models (VLMs) is typically performed at the image or instance level, making it difficult to precisely remove target knowledge without affecting unrelated semantics. This issue is especially pronounced since a single image often contains multiple entangled concepts, including both target concepts to be forgotten and contextual information that should be preserved. In this paper, we propose an interpretable concept-level unlearning framework for VLMs, which constructs a compact task-specific concept vocabulary from the forgetting set using a multimodal large language model. In addition to modality alignment, visual representations are decomposed into sparse, nonnegative combinations of semantic concepts, providing an explicit interface for fine-grained knowledge manipulation. Based on this decomposition, our method formulates unlearning as concept-level optimization, where target concepts are selectively suppressed while intra-instance non-target semantics and global cross-modal knowledge are preserved. Extensive experiments across both in-domain and out-of-domain forgetting settings demonstrate that our method enables more comprehensive target forgetting, better preserves non-target knowledge within the same image, and maintains competitive model utility compared with existing VLM unlearning methods.
May 26, 2026cs.CV

On the Robustness of Machine Unlearning for Vision-Language Models

Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work, we present the first systematic survey and robustness analysis of VLM unlearning. We provide a comprehensive taxonomy and review of existing VLM unlearning methods, together with unified evaluations under multiple prompt settings. We then propose three attack paradigms to examine whether forgotten multimodal knowledge can be reactivated through contextual prompting or downstream retraining. Extensive experiments show that many existing methods remain vulnerable under these attacks, indicating that current approaches often hide rather than fully remove target knowledge. Our study provides new insights into the robustness and limitations of current VLM unlearning methods and highlights the need for more reliable multimodal unlearning strategies. Code is available at https://github.com/XMUDeepLIT/VLM-UnL-Attack.
May 8, 2026cs.CV

Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models

Vision-language models (VLMs) raise growing concerns about privacy, copyright, and bias, motivating machine unlearning to remove sensitive knowledge. However, existing methods primarily fine-tune the language decoder, leading to superficial forgetting that fails to erase underlying visual representations and often introduces object hallucination. We propose HFRU, a reinforcement unlearning framework that operates on the vision encoder for deep semantic removal. Our two-stage approach combines alignment disruption with GRPO-based optimization using a composite reward, including an abstraction reward that encourages semantically valid substitutions and mitigates hallucinations. Experiments on object recognition and face identity tasks show that HFRU achieves over 98% forgetting and retention performance, while introducing negligible object hallucination, significantly outperforming prior methods.Our code and implementation details are available at https://github.com/XMUDeepLIT/HFRU.