cs.LGDec 28, 2025

Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models

Authors: Mingyuan Zhang, Yue Bai, Yifan Wang, Yiyang Huang, Yun Fu

Organizations: College of Engineering, Northeastern University · Khoury College of Computer Science, Northeastern University

Abstract

Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tuning (PEFT) methods rely on additive modules, such as low-rank adapters, which may limit adaptation capacity. In this paper, we rethink VLM adaptation from a structural selection framework that adapts VLMs without modifying backbone weights, and we propose Mask Fine-Tuning (MFT). MFT learns masks that selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align pretrained representations with downstream objectives. Extensive experiments show that MFT provides an effective structural alternative to both FFT and PEFT, consistently achieving superior performance across multiple vision-language benchmarks without adding knowledge or altering the deployment architecture. Moreover, our analysis with MFT provides new insights into how pretrained VLMs reorganize their internal representational pathways during adaptation.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders

    Sep 14, 2026Qingtao Xia, Jiahua Bao, Siyao Cheng +1Parameter-Efficient Fine-Tuning MethodsVision-Language Model Adaptation

  2. Boosting Large Language Models with Mask Fine-Tuning

    Mar 27, 2025Mingyuan Zhang, Yue Bai, Huan Wang +4Large Language Model Fine-Tuning

  3. Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models

    May 19, 2026Jia-Wei Hai, Yijun Wang, Xiu-Shen WeiVision-Language Model AdaptationTest-Time Adaptation