Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tuning (PEFT) methods rely on additive modules, such as low-rank adapters, which may limit adaptation capacity. In this paper, we rethink VLM adaptation from a structural selection framework that adapts VLMs without modifying backbone weights, and we propose Mask Fine-Tuning (MFT). MFT learns masks that selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align pretrained representations with downstream objectives. Extensive experiments show that MFT provides an effective structural alternative to both FFT and PEFT, consistently achieving superior performance across multiple vision-language benchmarks without adding knowledge or altering the deployment architecture. Moreover, our analysis with MFT provides new insights into how pretrained VLMs reorganize their internal representational pathways during adaptation.
Figures & tables
Figure 1: Comparison of fine-tuning methods on a VLM (Qwen2.5-0.5B language model).
Figure 2: Overview of Mask Fine-Tuning (MFT) . All pretrained parameters remain frozen, while learnable scores S are optimized for the projector and selected linear layers of the language backbone. The scores are mapped to masks M=g(S) , which modulate the pretrained weights W through elementwise multiplication, yielding the effective weights W=W⊙M .
Figure 3: Hyperparameter sensitivity of S-MFT with respect to the initial score value ( Sinit ) and temperature ( T ) across four language backbones. Each surface plot shows the MMMU performance under different hyperparameter combinations. The red dashed line indicates the configuration selected for full training.
Table 4
Figure 4: Layer-wise analysis of S-MFT across four language backbones. Each curve shows the average performance when masks are applied only to a subset of layers, compared with full-layer training ( red dashed line ).
Figure 6
Figure 7: Loss landscape of the Qwen2.5-0.5B backbone.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset Name
Samples
Usage
Pretraining
LAION-CC-SBU-558K ( Liu et al., 2023b )
558,128
Training
Supervised Fine-Tuning (SFT)
LAION-CC-SBU-558K ( Liu et al., 2023b )
558,128
Training
COCO (train2017) ( Lin et al., 2015 )
118,287
Training
GQA ( Ainslie et al., 2023 )
113,018
Training
Appendix
Table 5: Overview of datasets used across the pretraining, SFT, and evaluation stages. Training set sizes are denoted by the number of images, while evaluation sets are denoted by the number of query instances (QAs).
Figure 8: Performance comparison of S-MFT, FFT, and LoRA under varying training data ratios across four language backbones. Each curve shows the MMMU benchmark performance, with shaded regions indicating the standard deviation. S-MFT consistently outperforms FFT and LoRA across varying data ratios, with particularly pronounced advantages in low-data regimes.
We propose a pre-fine-tuning probing method for Parameter-Efficient Fine-Tuning (PEFT) layer selection, aiming to obtain more stable and higher gains with fewer trainable parameters when adapting large vision--language models (VLMs). Unlike the common practice of applying LoRA and other adapters to all layers at once---where layer selection often relies on heuristic rules---we focus on the vision encoder and directly evaluate the "adaptability'' of each Transformer layer. Specifically, we characterize each layer from two perspectives: (i) the statistical properties of its Q/K/V projection weights (e.g., norms and condition numbers); (ii) robustness under controlled parameter perturbations. We then systematically compare these indicators with the downstream performance gains brought by applying PEFT to a single layer. Across experiments covering seven benchmarks and five PEFT variants, we observe a consistent correlation: layers (or matrices) with larger weight norms and higher condition numbers are usually more robust to perturbations and are more likely to yield larger fine-tuning gains. These results show that distribution-statistics analysis and perturbation tests before fine-tuning can provide practical signals for adaptation-layer selection, thereby maintaining or improving performance while reducing trainable parameters.
Qingtao Xia, Jiahua Bao, Siyao Cheng +1
Faculty of Computing Harbin Institute of Technology Harbin, China
The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performance. In this work, we introduce Mask Fine-Tuning (MFT), a novel LLM fine-tuning paradigm demonstrating that carefully breaking the model's structural integrity can surprisingly improve performance without updating model weights. MFT learns and applies binary masks to well-optimized models, using the standard LLM fine-tuning objective as supervision. Based on fully fine-tuned models, MFT uses the same fine-tuning datasets to achieve consistent performance gains across domains and backbones (e.g., an average gain of 2.70/4.15 on IFEval with LLaMA2-7B/3.1-8B). Detailed ablation studies and analyses examine the proposed MFT from different perspectives, including the sparse ratio and the loss surface. Additionally, when deployed on well-trained models, MFT is compatible with other LLM optimization procedures to improve overall model performance. Furthermore, this study extends the masking operation beyond its conventional use in network pruning for model compression to encompass a broader range of model capabilities.
Mingyuan Zhang, Yue Bai, Huan Wang +4
College of Engineering, Northeastern University · Khoury College of Computer Science, Northeastern University
Vision-Language Models (VLMs), such as CLIP, have achieved significant zero-shot performance on downstream tasks with various fine-tuning adaptation methods. However, recent studies have proven that adversarial attacks can significantly degrade the inference ability of VLMs, posing substantial risks to their practical applications. Prevalent test-time adaptation methods typically rely on multi-view augmentation to implement various fine-tuning strategies, which struggle to identify semantic information and are prone to destroying discriminative regions in fine-grained scenarios. To address these limitations, we propose Attention-Guided Test-Time Prompt Tuning (A-TPT), a semantics-preserving method designed for test-time adaptation. We first refine the gradient attention rollout mechanism to identify semantically meaningful regions surviving under adversarial attacks. Furthermore, we leverage them to guide the spatially varying augmentation intensities and multi-view ensemble for prompt tuning and inference. Extensive experiments demonstrate that A-TPT outperforms existing test-time adaptation methods on both adversarial and clean data. Codes are available at https://github.com/SEU-VIPGroup/A-TPT .
Jia-Wei Hai, Yijun Wang, Xiu-Shen Wei
School of Computer Science and Engineering, and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, China. · School of Intelligence Science and Engineering, Southeast University, China.