Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tuning (PEFT) methods rely on additive modules, such as low-rank adapters, which may limit adaptation capacity. In this paper, we rethink VLM adaptation from a structural selection framework that adapts VLMs without modifying backbone weights, and we propose Mask Fine-Tuning (MFT). MFT learns masks that selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align pretrained representations with downstream objectives. Extensive experiments show that MFT provides an effective structural alternative to both FFT and PEFT, consistently achieving superior performance across multiple vision-language benchmarks without adding knowledge or altering the deployment architecture. Moreover, our analysis with MFT provides new insights into how pretrained VLMs reorganize their internal representational pathways during adaptation.
Figures & tables
Figure 1: Comparison of fine-tuning methods on a VLM (Qwen2.5-0.5B language model).
Figure 2: Overview of Mask Fine-Tuning (MFT) . All pretrained parameters remain frozen, while learnable scores S are optimized for the projector and selected linear layers of the language backbone. The scores are mapped to masks M=g(S) , which modulate the pretrained weights W through elementwise multiplication, yielding the effective weights W=W⊙M .
Figure 3: Hyperparameter sensitivity of S-MFT with respect to the initial score value ( Sinit ) and temperature ( T ) across four language backbones. Each surface plot shows the MMMU performance under different hyperparameter combinations. The red dashed line indicates the configuration selected for full training.
Table 4
Figure 4: Layer-wise analysis of S-MFT across four language backbones. Each curve shows the average performance when masks are applied only to a subset of layers, compared with full-layer training ( red dashed line ).
Figure 6
Figure 7: Loss landscape of the Qwen2.5-0.5B backbone.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset Name
Samples
Usage
Pretraining
LAION-CC-SBU-558K ( Liu et al., 2023b )
558,128
Training
Supervised Fine-Tuning (SFT)
LAION-CC-SBU-558K ( Liu et al., 2023b )
558,128
Training
COCO (train2017) ( Lin et al., 2015 )
118,287
Training
GQA ( Ainslie et al., 2023 )
113,018
Training
Appendix
Table 5: Overview of datasets used across the pretraining, SFT, and evaluation stages. Training set sizes are denoted by the number of images, while evaluation sets are denoted by the number of query instances (QAs).
Figure 8: Performance comparison of S-MFT, FFT, and LoRA under varying training data ratios across four language backbones. Each curve shows the MMMU benchmark performance, with shaded regions indicating the standard deviation. S-MFT consistently outperforms FFT and LoRA across varying data ratios, with particularly pronounced advantages in low-data regimes.
School of Computer Science and Engineering, and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, China. · School of Intelligence Science and Engineering, Southeast University, China.