cs.CRSep 29, 2026

Selective Channel Restoration for Backdoored Vision-Language Models

Authors: Shuming Liu, Zhifang Zhang, Suqin Yuan, Khin Mi Mi Aung, Zhuoyi Lin, Lei Feng

Organizations: Southeast University · The University of Queensland · University of Sydney · A*STAR

Abstract

Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these limitations, we propose Perturb-Select-Restore (PSR), a post-training defense that performs sparse updates to the projection interface and introduces no additional computation during inference. We reveal that backdoored VLM projectors are substantially more sensitive to bounded perturbations than clean VLM projectors, a phenomenon we term projection fragility. Building on this finding, PSR identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values. Experiments across multiple tasks show that PSR reduces attack success rates to near zero while preserving clean-task performance.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks

    Jun 1, 2026Ivan Sabolić, Marin Oršić, Josip Šarić +1

  2. CBV: Clean-label Backdoor Attacks on Vision Language Models via Diffusion Models

    May 4, 2026Ji Guo, Xiaolong Qin, Cencen Liu +3Clean Label Backdoor AttackDiffusion Models

  3. Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

    Jul 28, 2026Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cinà +3Vision-Language Foundation ModelsLarge Language Model Backbones