The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performance. In this work, we introduce Mask Fine-Tuning (MFT), a novel LLM fine-tuning paradigm demonstrating that carefully breaking the model's structural integrity can surprisingly improve performance without updating model weights. MFT learns and applies binary masks to well-optimized models, using the standard LLM fine-tuning objective as supervision. Based on fully fine-tuned models, MFT uses the same fine-tuning datasets to achieve consistent performance gains across domains and backbones (e.g., an average gain of 2.70/4.15 on IFEval with LLaMA2-7B/3.1-8B). Detailed ablation studies and analyses examine the proposed MFT from different perspectives, including the sparse ratio and the loss surface. Additionally, when deployed on well-trained models, MFT is compatible with other LLM optimization procedures to improve overall model performance. Furthermore, this study extends the masking operation beyond its conventional use in network pruning for model compression to encompass a broader range of model capabilities.
Figures & tables
Figure 1: The visualization of the performance trend across different fine-tuning strategies, including FFT (blue line), LoRA (green line) ( Hu et al., 2022 ) , and our MFT (red line). We also add random (orange dashed) and L1 (gray dashed) masks for comparison. We use three settings across LLaMA and Qwen backbones on GSM8K, HumanEval, and IF-Eval for math, coding, and instruction domains, respectively. The x-axis is training steps starting from the pre-trained backbone. The y-axis is evaluation performance. MFT (red line) starts from the best FFT model (yellow star) and improves upon the upper bound, whereas continued FFT leads to overfitting. It also outperforms LoRA fine-tuning and two vanilla mask baselines.
Figure 2: Typical LLM training comprises pre-training and fine-tuning to build capacity and domain knowledge, with the structure always being the entire model. We ask whether this integrity is necessary and propose that MFT generally outperforms models with sufficient FFT, naturally upgrading the classic pipeline by following the typical protocol to further refine well-optimized LLMs.
Figure 3: Visualization of the ablation study of the local MFT strategy. It uses LLaMA2-7B and Qwen3-1.7B backbones, covering math, coding, and instruction domains. In each figure, we conduct MFT ablations with a 10% masking ratio on a domain-specific FFT model (black dashed line). We swap the ablation between two local granularities, 8/7-layer (purple) and 4-layer (orange), moving from shallow to deep layers using a 10% fine-tuning set. This ablation is intended for quick intuition, with limited performance gains from MFT on a 10% subset. Based on this trend, we deploy MFT with complete fine-tuning sets, achieving more improvements.
Method
Math
Coding
Instruction Following
GSM8K
Math
HumanEval
HumanEval+
IF-Eval
Alpaca-Eval
LLaMA2-7B
Pre-Trained Model
15.2
2.5
25.8
22.4
34.3
0.5
Best LoRA
34.8±0.34
4.7±0.26
28.7±0.56
23.8±0.48
35.7±0.68
1.2±0.16
Continued LoRA
33.6±0.42
4.5±0.23
28.2±0.49
23.2±0.43
32.5±0.84
1.1±0.12
Best FFT
46.8±0.22
6.6±0.21
29.4±0.36
25.1±0.34
41.2±0.62
1.9±0.18
Continued FFT
45.0±0.28 ↓ 1.8
5.5±0.23 ↓ 1.1
27.9±0.42 ↓ 1.5
23.6±0.38 ↓ 1.5
37.8±0.68 ↓ 3.4
2.0±0.21 ↑ 0.1
Table 1: Performance comparison of domain-specific settings. Each part has three blocks containing 1) common LoRA and its continued variant (green), 2) FFT serving as an upper bound and its continued variant (blue), and 3) our MFT (red) with the other two vanilla masking baselines. The pre-trained performance serves as a lower bound.
Figure 4: Training cost comparisons of GPU memory, token usage, and training time. We use the coding domain as an example to compare MFT with other FFT and LoRA baselines.
Figure 5: Data ratio ablation. We use math and instruction domains on LLaMA2 and LLaMA3.1. Compared with Best FFT (red dashed line), MFT (purple) consistently improves performance on the full dataset but may still yield a promising gain with less data.
Figure 6: Masking ratio ablation. We use coding and instruction domains on LLaMA2 and LLaMA3.1. We observe that the original 10% ratio works well for coding but not for instruction, indicating that the masking ratio matters and that MFT has more potential.
Figure 7: The loss landscape visualization on math and instruction-following domains using LLaMA2-7B. Such visualization indicates that the proposed MFT further refines the Best FFT model, improving optimization and generalization.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Train / Test
Dataset Name
# of Samples
Math
Train
Tulu 3 Persona MATH ( Lambert et al., 2025 )
149,960
Tulu 3 Persona GSM ( Lambert et al., 2025 )
49,980
Tulu 3 Persona Algebra ( Lambert et al., 2025 )
20,000
MetaMathQA ( Yu et al., 2024 )
395,000
NuminaMath-TIR ( LI et al., 2024 )
64,312
Test
GSM8K ( Cobbe et al., 2021 )
1,320
Appendix
Table 2: The information on the training and test datasets used in our experiments.
Method
Math
Coding
Instruction Following
GSM8K
Math
HumanEval
HumanEval+
IF-Eval
Alpaca-Eval
Pre-trained Model
15.2
2.5
25.8
22.4
34.3
0.5
Best FFT [Math]
46.9
6.7
17.7
15.9
29.3
1.5
Best FFT [Coding]
14.6
2.8
29.3
25.0
8.3
1.4
Best FFT [IF]
25.0
2.4
16.5
13.4
41.4
1.7
Continued FFT [Math]
45.2
5.7
19.5
17.1
33.2
1.5
Appendix
Table 3: Cross-domain evaluations on LLaMA2-7B under domain-specific FFT setting.
Method
Math
Coding
Instruction Following
GSM8K
Math
HumanEval
HumanEval+
IF-Eval
Alpaca-Eval
Pre-trained Model
55.9
14.6
42.1
37.8
17.6
9.0
Best FFT [Math]
77.0
24.3
35.4
29.9
31.5
0.6
Best FFT [Coding]
60.7
13.8
51.2
45.1
5.3
0.9
Best FFT [IF]
63.0
10.1
50.0
44.5
59.8
12.0
Continued FFT [Math]
74.6
24.0
33.5
31.1
32.2
0.7
Appendix
Table 4: Cross-domain evaluations on LLaMA3.1-8B under domain-specific FFT setting.
Method
Math
Coding
Instruction Following
GSM8K
Math
HumanEval
HumanEval+
IF-Eval
Alpaca-Eval
Pre-Trained Model
15.2
2.5
25.8
22.4
34.3
0.5
Mixed Domain
Best LoRA
40.8±0.30
6.1±0.28
22.6±0.55
18.3±0.44
37.3±0.81
0.7±0.11
Continued LoRA
31.9±0.24
4.0±0.21
20.1±0.42
17.1±0.36
31.5±0.92
0.8±0.09
Best FFT
45.5±0.18
8.1±0.26
29.7±0.43
26.7±0.38
43.6±0.68
1.0±0.15
Continued FFT
44.1±0.22 ↓ 1.4
7.5±0.27 ↓ 0.6
21.1±0.58 ↓ 8.6
18.1±0.43 ↓ 8.6
38.6±0.74 ↓ 5.0
1.2±0.16 ↑ 0.2
Appendix
Table 5: Performance comparison on LLaMA2-7B and LLaMA3.1-8B in multi-domain settings. Each part has three blocks containing 1) common LoRA and its continued variant (green), 2) FFT serving as an upper bound and its continued variant (blue), and 3) our MFT (red) with the other two vanilla masking baselines. The pre-trained performance serves as a lower bound.
Method
LLaMA2-7B
LLaMA3.1-8B
GSM8K
Math
GSM8K
Math
Pre-Trained Model
15.2
2.5
55.9
14.6
Best FFT
46.9
6.7
77.0
24.3
Continued FFT
45.2 ↓ 1.7
5.7 ↓ 1.0
74.6 ↓ 2.4
24.0 ↓ 0.3
Global MFT w/ Best FFT (Ours)
49.0 ↑ 2.1
7.1 ↑ 0.4
74.1 ↓ 2.9
21.8 ↓ 2.5
Appendix
Table 6: Initial exploration of global MFT on the math domain using LLaMA2-7B and LLaMA3.1-8B backbones. We observe better performance on LLaMA2-7B, especially on GSM8K, but lower performance on LLaMA3.1-8B.
Method
GQA
MMMU
POPE
MME
SQA
TextVQA
LLaVA (CLIP + Qwen2.5-0.5B)
Pre-Trained Model
5.7
22.6
49.9
93.7
3.2
9.0
Best FFT
55.7
34.2
85.3
1197.9
57.6
44.2
MFT (Ours)
57.5 ↑ 1.8
35.0 ↑ 0.8
86.3 ↑ 1.0
1250.2 ↑ 52.3
58.8 ↑ 1.2
46.1 ↑ 1.9
LLaVA (SigLIP + Gemma-2B)
Pre-Trained Model
2.6
26.4
56.6
586.2
2.8
5.6
Appendix
Table 7: Performance of MFT compared to Best FFT on multi-modal benchmarks across two LLaVA backbones. MFT consistently outperforms Best FFT on all evaluated metrics.
Method
MBPP
Zero-Shot
46.8
Best FFT
50.5
+ Dropout
50.7
+ Stronger Weight Decay
50.5
+ MFT (Ours)
51.2
Appendix
Table 8: Comparison of MFT with the best FFT and common regularization techniques.
Figure 8: The local MFT ablation results of the larger layer-wise group (16-layer).
LLaMA2-7B
Math
Coding
Instruction
FFT loss
0.101
0.098
0.125
MFT loss
0.085
0.054
0.060
Appendix
Table 9: Exemplar training loss statistics of LLaMA2-7B models after training are stable.
Domain
Best FFT
Continued FFT
Best MFT
Complete MFT
Math
2.4 epochs (5.90e8)
4.0 epochs (9.83e8)
4.0 epochs (9.83e8)
4.4 epochs (1.08e9)
Coding
2.7 epochs (3.32e8)
4.0 epochs (4.92e8)
3.6 epochs (4.43e8)
4.7 epochs (5.78e8)
Instruction Following
2.4 epochs (5.51e8)
4.0 epochs (9.18e8)
3.2 epochs (7.35e8)
4.4 epochs (1.01e9)
Appendix
Table 10: Training epochs and corresponding number of used tokens (in brackets) on LLaMA2-7B under domain‐specific FFT setting.
Domain
Best FFT
Continued FFT
Best MFT
Complete MFT
Math
2.0 epochs (2.67e9)
4.0 epochs (5.34e9)
3.7 epochs (3.09e9)
4.0 epochs (3.16e9)
Coding
2.0 epochs (2.67e9)
4.0 epochs (5.34e9)
3.3 epochs (2.83e9)
4.0 epochs (2.92e9)
Instruction Following
2.0 epochs (2.67e9)
4.0 epochs (5.34e9)
4.0 epochs (3.13e9)
4.0 epochs (3.13e9)
Appendix
Table 11: Training epochs and corresponding number of used tokens (in brackets) on LLaMA2-7B under mixed-up FFT setting.
Figure 9: Training cost comparisons of GPU memory, used token number, and training time on math and instruction-following domains using LLaMA2-7B and LLaMA3.1-8B.
Figure 10: Masking ratio ablation study on the math domain using LLaMA2-7B and LLaMA3.1-8B.
Figure 11: Data ratio ablation study on coding domain using LLaMA2-7B and LLaMA3.1-8B.
Figure 12: The loss landscape visualization on the coding domain using LLaMA2-7B.
Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tuning (PEFT) methods rely on additive modules, such as low-rank adapters, which may limit adaptation capacity. In this paper, we rethink VLM adaptation from a structural selection framework that adapts VLMs without modifying backbone weights, and we propose Mask Fine-Tuning (MFT). MFT learns masks that selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align pretrained representations with downstream objectives. Extensive experiments show that MFT provides an effective structural alternative to both FFT and PEFT, consistently achieving superior performance across multiple vision-language benchmarks without adding knowledge or altering the deployment architecture. Moreover, our analysis with MFT provides new insights into how pretrained VLMs reorganize their internal representational pathways during adaptation.
Mingyuan Zhang, Yue Bai, Yifan Wang +2
College of Engineering, Northeastern University · Khoury College of Computer Science, Northeastern University
Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater representational plasticity, Low-Rank Adaptation (LoRA) can match or surpass FFT performance while constraining updates to a low-rank space and potentially benefiting from additional regularization. Through empirical evaluation across diverse tasks (SQL, Medical QA, and Counterfactual Knowledge) and varying language models (Gemma-3-1B, Qwen2.5-1.5B, and Qwen2.5-3B), we observe both trends and find that the better static architecture depends on the task and model. Spectral and truncation analyses further show that endpoint compressibility alone does not explain these task differences, suggesting task-score sensitivity and constrained optimization trajectories as possible explanations. To address this challenge, we propose a Mixture of LoRA and Full (MoLF) Fine-Tuning, a unified framework that enables continuous navigation between both training regimes. MoLF dynamically routes updates between FFT and LoRA at the optimizer level to ensure that exact gradient signals are available to both experts throughout training, while only selected experts update their weights. For memory-constrained environments, we also introduce MoLF-Efficient, which freezes base weights and only routes updates among a pair of LoRA experts of potentially varying rank. Our evaluations show that MoLF either improves on or stays within 1.5 percentage points of the better of FFT and LoRA across the nine tested settings, while MoLF-Efficient outperforms both AdaLoRA and AdaMix in eight of nine settings, with gains over the stronger baseline of up to 11.70 percentage points on Fact, 3.13 on Med, and 2.98 on SQL.
Haozhan Tang, Xiuqi Zhu, Xinyin Zhang +3
Carnegie Mellon University · Tsinghua University · Infinigence AI
Large language models (LLMs) remain expensive to fine-tune because full-parameter updates require substantial memory, compute, and per-task storage. We study whether saliency signals originally developed for pruning can be reused to choose where a model should adapt. We propose Super, a sparse parameter-efficient fine-tuning (PEFT) method that fixes a small trainable support using a Wanda-style activation-weighted magnitude score [Sun et al., 2023] computed from a calibration pass. We then introduce Supra, a hybrid adapter that combines this sparse update with LoRA while preserving a matched trainable-parameter budget through a simple budget-splitting rule. In single-seed Math17K arithmetic experiments on Llama-3.2-1B and Meta-Llama-3-8B, the best Super/Supra variants achieve the highest average accuracy among the tested schedule-selected adapter configurations. We also include a PaFi-style magnitude-only support as a closest training-free sparse baseline and find that low-score supports under both magnitude and Wanda-style orderings can be effective. These results suggest that simple pruning-inspired orderings can provide useful fixed sparse supports for PEFT, especially when combined with low-rank adapters.