Open-MMUnlearning: Unifying Methods and Evaluation for MLLM Unlearning
Organizations: Institute of Automation, Chinese Academy of Sciences
Abstract
As multimodal large language models (MLLMs) become more capable and widely deployed, concerns about privacy and safety have become increasingly pressing. Machine unlearning offers one approach to addressing these concerns by removing designated information from trained models while preserving unrelated capabilities. However, fragmented implementations and evaluation protocols, incomplete robustness testing, and limited understanding of metric reliability make progress in MLLM unlearning difficult to assess systematically. We introduce Open-MMUnlearning, an open-source, extensible framework that integrates target-model preparation, multimodal data processing, unlearning, and evaluation through shared interfaces and structured configurations. The framework supports five benchmarks spanning privacy, safety, and copyright, eight MLLMs from four model families, and twelve unlearning methods. Its evaluation suite jointly assesses forgetting effectiveness, retained utility, and robustness to model interventions, adversarial inputs, and membership inference attacks. Using a common evaluation protocol, we compare ten representative unlearning methods. In this comparison, GD and MIP-Editor tie for the highest overall score: GD achieves the highest Forget Quality, while MIP-Editor preserves more Model Utility. We further introduce a metric meta-evaluation protocol that tests faithfulness using models with controlled exposure to target knowledge and robustness under quantization and relearning. Among the thirteen evaluated metrics, BLEU achieves the highest aggregate reliability score. KS-Test attains the highest faithfulness AUC but performs less well on robustness. Together, the framework and these findings support reproducible comparison of MLLM unlearning methods and systematic assessment of evaluation reliability.
Figures & tables
| Component | Variants | |
| Models | LLaVA -1.5 ( Liu et al., 2023 ) Qwen -2.5-VL ( Bai et al., 2025 ) LLaVA -1.6 ( Liu et al., 2024a ) Gemma -3 ( Team et al., 2025 ) | |
| Methods | General Unlearning Methods ( Maini et al., 2024 ; Li et al., 2024b ) ( Zhang et al., 2024a ; Dong et al., 2025 ) MMUnlearner ( Huo et al., 2025 ) MANU ( Liu et al., 2025b ) MIP-Editor ( Li et al., 2026 ) SMFA ( Zeng et al., 2025 ) VGID ( Chen et al., 2026b ) | |
| Datasets | MLLMU ( Liu et al., 2025a ) FIUBench ( Ma et al., 2025 ) CLEAR ( Dontsov et al., 2025 ) SafeEraser ( Chen et al., 2025 ) CoVUBench ( Kwon et al., 2026 ) | |
| Metrics | Forget | Truth Ratio ( Ma et al., 2025 ; Dontsov et al., 2025 ) KS-Test ( Ma et al., 2025 ; Dontsov et al., 2025 ) JS Distance ( Dontsov et al., 2025 ) Answer Probability ( Ma et al., 2025 ) ROUGE-L ( Liu et al., 2025a ) ASR ( Chen et al., 2025 ) |
| Utility | Model Utility ( Ma et al., 2025 ; Dontsov et al., 2025 ) Fluency ( Kwon et al., 2026 ) Specificity ( Kwon et al., 2026 ) Generality ( Kwon et al., 2026 ) SARR ( Chen et al., 2025 ) LM-Eval ( Liu et al., 2025a ; Kwon et al., 2026 ) | |
| Robust | Relearning ( Hu et al., 2025 ; Zheng et al., 2025 ) Quantization ( Zhang et al., 2025c ) Probing ( Lynch et al., 2024 ) MIA ( Yeom et al., 2018 ; Carlini et al., 2021 ; Wang et al., 2024 ; Shi et al., 2024 ) Jailbreak ( Wei et al., 2023 ; Ma et al., 2024 ) FigStep ( Gong et al., 2025 ) SUA ( Zhang et al., 2025b ) Image Rephrasing ( Patil et al., 2025 ) |
| Methods | Agg. | Forget Quality , | Model Utility , | Robustness | |||
| Agg. | Model | Input | Output | ||||
| Reference Models | |||||||
| Finetune | 0.622 | 0.687 | 0.421 | 0.759 | 0.835 | 0.631 | 0.812 |
| Retain-Only | 0.638 | 0.696 | 0.444 | 0.776 | 0.839 | 0.688 | 0.799 |
| General Unlearning Methods | |||||||
| GA ( Maini et al., 2024 ) | 0.574 | 0.774 | 0.066 | 0.881 | 0.996 | 0.971 | 0.677 |
| Metrics | Agg. | Faithful. | Robustness | ||
| Agg. | Quant. | Relearn | |||
| Cloze | 0.766 | 0.626 | 0.986 | 0.995 | 0.978 |
| Classification | 0.740 | 0.604 | 0.954 | 0.978 | 0.930 |
| Truth Ratio | 0.896 | 0.814 | 0.996 | 0.999 | 0.992 |
| Probability | 0.887 | 0.924 | 0.853 | 0.928 | 0.789 |
| KS-Test | 0.888 | 0.991 | 0.804 | 0.900 | 0.727 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Variants | |
| Models | LLaVA -1.5 ( Liu et al., 2023 ) Qwen -2.5-VL ( Bai et al., 2025 ) LLaVA -1.6 ( Liu et al., 2024a ) Gemma -3 ( Team et al., 2025 ) | |
| Methods | GradAscent GradDiff KL ( Maini et al., 2024 ) IdkNLL ( Maini et al., 2024 ) RMU ( Li et al., 2024b ) NPO ( Zhang et al., 2024a ) UNDIAL ( Dong et al., 2025 ) MMUnlearner ( Huo et al., 2025 ) MANU ( Liu et al., 2025b ) MIP-Editor ( Li et al., 2026 ) SMFA ( Zeng et al., 2025 ) VGID ( Chen et al., 2026b ) | |
| Datasets | MLLMU ( Liu et al., 2025a ) FIUBench ( Ma et al., 2025 ) CLEAR ( Dontsov et al., 2025 ) SafeEraser ( Chen et al., 2025 ) CoVUBench ( Kwon et al., 2026 ) | |
| Metrics | Forget | Truth Ratio ( Ma et al., 2025 ; Dontsov et al., 2025 ) Exact Match / APE ( Ma et al., 2025 ) Answer Probability ( Ma et al., 2025 ) Fill-in-the-Blank Accuracy ( Liu et al., 2025a ) Classification Accuracy ( Liu et al., 2025a ) ROUGE-1/2/L ( Liu et al., 2025a ) BLEU ( Liu et al., 2025a ) KS-Test ( Ma et al., 2025 ; Dontsov et al., 2025 ) JS Distance ( Dontsov et al., 2025 ) Forget Quality ( Dontsov et al., 2025 ) Keyword Recall ( Kwon et al., 2026 ) Semantic Dissimilarity ( Kwon et al., 2026 ) Efficacy ( Kwon et al., 2026 ) Divergence ( Kwon et al., 2026 ) Attack Success Rate (ASR) ( Chen et al., 2025 ) Refusal Rate (RR) ( Chen et al., 2025 ) |
| Utility | Model Utility ( Ma et al., 2025 ; Dontsov et al., 2025 ) GPT-Eval ( Ma et al., 2025 ; Chen et al., 2025 ) Fluency ( Kwon et al., 2026 ) Specificity ( Kwon et al., 2026 ) Generality ( Kwon et al., 2026 ) LM-Eval ( Liu et al., 2025a ; Kwon et al., 2026 ) SARR ( Chen et al., 2025 ) | |
| Robust | Relearning ( Hu et al., 2025 ; Zheng et al., 2025 ) Quantization ( Zhang et al., 2025c ) Probing ( Lynch et al., 2024 ) Cross-Modal Leakage ( Wang et al., 2026a ) Pure-Text Jailbreak ( Zou et al., 2023 ; Wei et al., 2023 ; Ma et al., 2024 ) FigStep ( Gong et al., 2025 ) Image Rephrasing ( Patil et al., 2025 ) SUA ( Zhang et al., 2025b ) LOSS ( Yeom et al., 2018 ) ZLIB ( Carlini et al., 2021 ) GradNorm ( Wang et al., 2024 ) Min-K% ( Shi et al., 2024 ) Min-K%++ ( Zhang et al., 2025a ) |
| Methods | Forget Set | Test Set | ||||
| Cloze | Cls. | Gene. | Cloze | Cls. | Gene. | |
| Reference Models | ||||||
| Finetune | 13.54 | 38.78 | 0.242 | 11.46 | 27.35 | 0.149 |
| Retain-Only | 12.50 | 36.73 | 0.207 | 13.54 | 40.41 | 0.146 |
| General Unlearning Methods | ||||||
| GA ( Maini et al., 2024 ) | 0.00 | 0.00 | 0.003 | 0.00 | 0.00 | 0.003 |
| Methods | Truth Ratio | Prob. | KS-Test | JS Dist. |
| Reference Models | ||||
| Finetune | 0.570 | 0.499 | 0.140 | 0.281 |
| Retain-Only | 0.452 | 0.254 | 0.150 | 0.337 |
| General Unlearning Methods | ||||
| GA ( Maini et al., 2024 ) | 0.393 | 0.000 | 0.330 | 0.190 |
| GD ( Maini et al., 2024 ) | 0.432 | 0.168 | 0.130 | 0.130 |
| Methods | Retain-Shared | Retain-Celebrity | ||||
| Cloze | Cls. | Gene. | Cloze | Cls. | Gene. | |
| Reference Models | ||||||
| Finetune | 8.04 | 36.74 | 0.231 | 12.75 | 29.90 | 0.171 |
| Retain-Only | 16.52 | 41.47 | 0.261 | 16.34 | 50.91 | 0.169 |
| General Unlearning Methods | ||||||
| GA ( Maini et al., 2024 ) | 0.11 | 0.00 | 0.003 | 0.00 | 0.00 | 0.000 |
| Methods | MMBench | MM-Vet | POPE | SQA | VizWiz | GQA | VQAv2 |
| Reference Models | |||||||
| Finetune | 80.00 | 16.00 | 82.00 | 56.00 | 2.60 | 52.00 | 79.20 |
| Retain-Only | 80.00 | 24.00 | 78.00 | 62.00 | 8.40 | 44.00 | 79.20 |
| General Unlearning Methods | |||||||
| GA ( Maini et al., 2024 ) | 0.00 | 0.00 | 0.00 | 36.00 | 0.00 | 0.00 | 0.00 |
| GD ( Maini et al., 2024 ) | 76.00 | 16.00 | 82.00 | 56.00 | 5.20 | 56.00 | 81.20 |
| Methods | Quant. | Relearn. | Probe-L16 | Probe-L24 |
| Reference Models | ||||
| Finetune | 0.752 | 0.759 | 0.993 | 0.996 |
| Retain-Only | 0.768 | 0.760 | 0.984 | 0.993 |
| General Unlearning Methods | ||||
| GA ( Maini et al., 2024 ) | 0.999 | 0.997 | 0.992 | 0.991 |
| GD ( Maini et al., 2024 ) | 0.915 | 0.772 | 0.994 | 0.993 |
| Methods | Cross-Modal | FigStep , | Jailbreak , | Image | SUA , |
| Leakage , | Rephrasing , | ||||
| Reference Models | |||||
| Finetune | 0.574 | 0.860 | 0.740 | 0.688 | 0.292 |
| Retain-Only | 0.611 | 0.930 | 0.797 | 0.667 | 0.438 |
| General Unlearning Methods | |||||
| GA ( Maini et al., 2024 ) | 0.853 | 1.000 | 1.000 | 1.000 | 1.000 |
| Methods | LOSS | ZLIB | GradNorm | Min-K% | Min-K%++ |
| Reference Models | |||||
| Finetune | 0.840 | 0.748 | 0.536 | 0.969 | 0.965 |
| Retain-Only | 0.954 | 0.856 | 0.456 | 0.884 | 0.846 |
| General Unlearning Methods | |||||
| GA ( Maini et al., 2024 ) | 0.874 | 0.823 | 0.033 | 0.828 | 0.828 |
| GD ( Maini et al., 2024 ) | 0.872 | 0.969 | 0.000 | 0.665 | 0.590 |