Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs
Organizations: University of Rochester
Abstract
Vision-language models (VLMs) may need to forget visual concepts after deployment because of privacy, copyright, licensing, safety, or policy changes. Conventional machine unlearning modifies model parameters, which may be costly or inaccessible for API-only models. Prompt-based suppression offers a training-free alternative, but does it make a concept inaccessible or merely change the model's answer? We investigate this question in off-the-shelf VLMs. Our visually grounded, multi-probe evaluation first verifies that a model recognizes each concept from the image, then tests its recoverability through multiple-choice, short-answer, and indirect queries. Across objects, scenes, and identities, prompt suppression reduces short-answer recall for some concepts while leaving multiple-choice and indirect performance largely unchanged. Explicitly listing the concepts to suppress often increases short-answer recall, suggesting that the list itself cues the answer. Beyond prompting, decoding constraints, representation editing, and parameter updates can suppress particular responses while the same concept remains detectable through other queries. A model may stop naming a visual concept yet still identify or use it when queried differently.
Figures & tables
| Method | Target SA [LB,UB] | Target MCQ | Cross-probe gap, 95% CI | Retain SA | Retain MCQ |
|---|---|---|---|---|---|
| GA | |||||
| GD | |||||
| GA+KL | |||||
| NPO | |||||
| MMUnlearner |
| Semantic | ||||
|---|---|---|---|---|
| Method | Exact | Lexical | LB | UB |
| GA | 0.192 | 0.089 | 0.040 | 0.201 |
| GD | 0.031 | 0.031 | -0.022 | 0.024 |
| GA+KL | 0.090 | 0.067 | -0.019 | 0.034 |
| NPO | 0.210 | 0.084 | 0.040 | 0.215 |
| MMUnl. | 0.134 | 0.143 | 0.135 | 0.173 |