cs.CLMar 27, 2025

Boosting Large Language Models with Mask Fine-Tuning

Authors: Mingyuan Zhang, Yue Bai, Huan Wang, Yizhou Wang, Qihua Dong, Yitian Zhang, Yun Fu

Organizations: College of Engineering, Northeastern University · Khoury College of Computer Science, Northeastern University

Abstract

The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performance. In this work, we introduce Mask Fine-Tuning (MFT), a novel LLM fine-tuning paradigm demonstrating that carefully breaking the model's structural integrity can surprisingly improve performance without updating model weights. MFT learns and applies binary masks to well-optimized models, using the standard LLM fine-tuning objective as supervision. Based on fully fine-tuned models, MFT uses the same fine-tuning datasets to achieve consistent performance gains across domains and backbones (e.g., an average gain of 2.70/4.15 on IFEval with LLaMA2-7B/3.1-8B). Detailed ablation studies and analyses examine the proposed MFT from different perspectives, including the sparse ratio and the loss surface. Additionally, when deployed on well-trained models, MFT is compatible with other LLM optimization procedures to improve overall model performance. Furthermore, this study extends the masking operation beyond its conventional use in network pruning for model compression to encompass a broader range of model capabilities.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models

    Dec 28, 2025Mingyuan Zhang, Yue Bai, Yifan Wang +2Vision-Language Model AdaptationParameter-Efficient Fine-Tuning Methods

  2. Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation

    May 8, 2026Haozhan Tang, Xiuqi Zhu, Xinyin Zhang +3Training-Side Stage-Aware Low-Rank AdaptationLarge Language Model Adaptation