cs.LGSep 27, 2026

SMAT: Simple and Efficient Merge-Aware Training

Authors: Yanggan Gu, Yuanyi Wang, Zhen Li, Shuo Cai, Yuhang Liu, Junzhuo Li, Zihao Wang, Hongxia Yang

Organizations: The Hong Kong Polytechnic University (PolyU) · The Hong Kong University of Science and Technology (Guangzhou) · The Chinese University of Hong Kong · PolyU-Daya Bay Technology and Innovation Research Institute

Abstract

Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training

    May 26, 2026Wenjie Zhou, Bohan Wang, Hongtao Zhang +3Large Language Model PretrainingMerge

  2. Saliency-Aware Model Merging

    May 30, 2026Jungin Park, Jiyoung Lee, Kwanghoon SohnContinual Model MergingMulti--Task Learning