Organizations: The Hong Kong Polytechnic University (PolyU) · The Hong Kong University of Science and Technology (Guangzhou) · The Chinese University of Hong Kong · PolyU-Daya Bay Technology and Innovation Research Institute
Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.
Figures & tables
Figure 1: Overview of SMAT. (a) SMAT simulates merged parameter states during expert training through Scale, Mask, and Perturb. (b) This improves merged performance while retaining training speed close to standard fine-tuning.
Figure 2: Simulating merged updates. Scale reweights the task vector, Mask drops and rescales its coordinates, and Perturb adds uniform noise to simulate other experts’ updates.
Figure 3: Perturb alone on Llama-1B (AdamW, t=4 ). Five-merger means ± SD.
Figure 4: Efficient SMAT training. Each period of t=4 steps includes one SMAT update. Fused kernels and storage switching reduce parameter overhead.
Method \CT@row@color \CT@row@color
Expert ↑ \CT@row@color \CT@row@color
Merged Model Performance ↑ \CT@row@color \CT@row@color
Training cost ↓
\CT@row@color \CT@row@color
\CT@row@color \CT@row@color
WA
TA
TIES
DARE
DELLA
Avg. \CT@row@color \CT@row@color
Time ( × FT)
Mem. (GiB)
Llama-3.2-1B-Instruct
FT \CT@row@color \CT@row@color
54.43 \CT@row@color \CT@row@color
39.88 ↑ 0.0
42.64 ↑ 0.0
42.93 ↑ 0.0
43.50 ↑ 0.0
41.17 ↑ 0.0
42.02 ↑ 0.0 \CT@row@color \CT@row@color
1.000
20.0 ↑ 0.00%
ASAM \CT@row@color \CT@row@color
55.20 \CT@row@color \CT@row@color
38.05 ↓ 1.8
46.29 ↑ 3.6
45.48 ↑ 2.5
45.71 ↑ 2.2
44.56 ↑ 3.4
44.02 ↑ 2.0 \CT@row@color \CT@row@color
2.869
22.3 ↑ 11.5%
MergOPT \CT@row@color \CT@row@color
56.23 \CT@row@color \CT@row@color
39.67 ↓ 0.2
45.62 ↑ 3.0
45.12 ↑ 2.2
45.59 ↑ 2.1
44.54 ↑ 3.4
44.11 ↑ 2.1 \CT@row@color \CT@row@color
1.718
22.3 ↑ 11.5%
OrthoReg \CT@row@color \CT@row@color
54.44 \CT@row@color \CT@row@color
43.85 ↑ 4.0
45.77 ↑ 3.1
45.93 ↑ 3.0
45.16 ↑ 1.7
42.57 ↑ 1.4
44.66 ↑ 2.6 \CT@row@color \CT@row@color
5.693
28.2 ↑ 41.1%
Table 1: TRACE results with AdamW. Bold marks the best score for each model.
Method \CT@row@color \CT@row@color
Expert ↑ \CT@row@color \CT@row@color
Merged Model Performance ↑ \CT@row@color \CT@row@color
Training cost ↓
\CT@row@color \CT@row@color
\CT@row@color \CT@row@color
WA
TA
TIES
DARE
DELLA
Avg. \CT@row@color \CT@row@color
Time ( × FT)
Mem. (GiB)
CLIP ViT-B/32
FT \CT@row@color \CT@row@color
89.26 \CT@row@color \CT@row@color
67.22 ↑ 0.0
71.13 ↑ 0.0
73.05 ↑ 0.0
71.21 ↑ 0.0
71.75 ↑ 0.0
70.87 ↑ 0.0 \CT@row@color \CT@row@color
1.000
5.8 ↑ 0.00%
ASAM \CT@row@color \CT@row@color
90.61 \CT@row@color \CT@row@color
67.57 ↑ 0.4
72.78 ↑ 1.6
75.05 ↑ 2.0
72.78 ↑ 1.6
73.68 ↑ 1.9
72.37 ↑ 1.5 \CT@row@color \CT@row@color
1.730
6.1 ↑ 5.67%
MergOPT \CT@row@color \CT@row@color
90.05 \CT@row@color \CT@row@color
68.13 ↑ 0.9
73.15 ↑ 2.0
74.82 ↑ 1.8
72.91 ↑ 1.7
73.84 ↑ 2.1
72.57 ↑ 1.7 \CT@row@color \CT@row@color
1.146
6.1 ↑ 5.67%
OrthoReg \CT@row@color \CT@row@color
89.36 \CT@row@color \CT@row@color
70.74 ↑ 3.5
72.82 ↑ 1.7
74.37 ↑ 1.3
72.85 ↑ 1.6
70.11 ↓ 1.6
72.18 ↑ 1.3 \CT@row@color \CT@row@color
1.125
6.5 ↑ 13.8%
Table 2: CLIP results with Adam. Bold marks the best score for each model.
Method \CT@row@color \CT@row@color
Expert ↑ \CT@row@color \CT@row@color
Avg. ↑ \CT@row@color \CT@row@color
Time ↓ ( × FT)
FT \CT@row@color \CT@row@color
55.31 \CT@row@color \CT@row@color
38.25 ↑ 0.0 \CT@row@color \CT@row@color
1.00
ASAM \CT@row@color \CT@row@color
54.69 \CT@row@color \CT@row@color
40.34 ↑ 2.1 \CT@row@color \CT@row@color
1.66
MergOPT \CT@row@color \CT@row@color
54.94 \CT@row@color \CT@row@color
40.22 ↑ 2.0 \CT@row@color \CT@row@color
1.25
OrthoReg \CT@row@color \CT@row@color
50.89 \CT@row@color \CT@row@color
43.41 ↑ 5.2 \CT@row@color \CT@row@color
2.87
SMAT \CT@row@color \CT@row@color
55.07 \CT@row@color \CT@row@color
44.15 ↑ 5.9 \CT@row@color \CT@row@color
1.00
Table 3: TRACE with Muon on Llama-1B. Bold marks the best score.
Figure 5: Llama-1B expert merging with AdamW: (a) varying expert count, (b) mixing MAT and FT experts, and (c) expert scores (open) versus mean merged scores (filled). Relative scores in (a,b) are normalized to FT experts.
Figure 6: Loss slices for six tasks using the experts in Table 1 . Columns are FT, MergOPT, and SMAT; stars mark independent experts. Each task shares one NLL color scale across methods; scales differ across tasks, and contours are spaced by 0.1.
Variant
Avg. ↑
Change
Full SMAT
45.77
–
w/o Scale
44.52
−1.25
w/o Mask
44.70
−1.07
w/o Perturb
43.67
−2.10
Table 4: Component ablation on Llama-1B with AdamW. Avg. averages five mergers; Change is relative to Full SMAT.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
ASAM radius ρ
MergOPT Laplace scale bMergOPT
OrthoReg λreg
Llama-1B
2
5×10−4
0.03
Llama-8B
0.25
2.5×10−4
0.01
ViT-B/32
0.25
2×10−4
0.1
ViT-L/14
0.125
5×10−5
0.05
Appendix
Table 5: Baseline training hyperparameters for the four backbones.
Method
Expert
Merged
Avg.
WA
TA
TIES
DARE
DELLA
FT
55.31
37.06
40.46
38.53
40.92
34.26
38.25
ASAM
54.69
39.50
41.35
41.65
41.89
37.32
40.34
MergOPT
54.94
38.07
42.52
41.90
42.02
36.59
40.22
OrthoReg
50.89
42.60
44.73
43.98
45.42
40.35
43.41
SMAT
55.07
42.77
46.27
46.44
45.72
39.55
44.15
Appendix
Table 6: Muon results on Llama-3.2-1B-Instruct. Avg. is the mean over five merger methods.
Llama-1B
Llama-8B
ViT-B/32
ViT-L/14
Method
Mean
Total
Mean
Total
Mean
Total
Mean
Total
FT
0.0687
1,382.1
0.3543
7,130.1
0.1255
4,015.3
1.4494
46,381.4
ASAM
0.1971
3,965.7
0.9481
19,079.0
0.2171
6,946.1
2.8991
92,771.3
MergOPT
0.1180
2,375.2
0.5469
11,005.0
0.1438
4,601.5
1.4862
47,559.9
OrthoReg
0.3910
7,868.5
4.1047
82,599.3
0.1411
4,515.7
1.5006
48,019.5
SMAT
0.0700
1,409.0
0.3594
7,231.2
0.1268
4,057.5
1.4516
46,451.0
Appendix
Table 7: Training time for complete expert sets. Mean is seconds per optimizer step; total is summed wall time in seconds.
Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematically studying how training duration of domain experts affects the quality of the merged model. We fine-tune experts on five domains (Math, Code, Instruction Following, Multilingual, and Safety) across three model sizes (Qwen 3.5 0.8B, 2B, and 4B), saving checkpoints from 25% to 500% of the optimal training steps and evaluating five merging methods at each duration. Our findings reveal a striking method-dependent pattern: simple averaging degrades sharply with overfitting, while sparsification-based methods achieve their best performance well past the validation optimum. We formalize this through bias-variance decomposition analysis, drawing a parallel to random forests where averaging benefits from high-variance individual learners. These results suggest that training duration and merging method should be chosen jointly rather than independently.
Model merging has emerged as a lightweight paradigm for enhancing Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. In this work, we analyze late-stage pre-training trajectories and uncover a \textbf{Rank-1 Subspace} phenomenon: while raw optimization steps oscillate violently, consecutive \emph{merged} checkpoints collapse onto a stable, approximately one-dimensional linear manifold. We theoretically ground this observation in a \emph{river-valley} landscape analysis: averaging acts as a geometric low-pass filter that dampens high-curvature noise to reveal the optimal descent direction. Capitalizing on this insight, we propose \textbf{Extra-Merge}, a training-free strategy that extrapolates along this subspace to minimize loss without additional gradient updates. Extensive experiments across GPT-2 and LLaMA families (124M to 2B) demonstrate that Extra-Merge consistently outperforms standard merging baselines. Notably, it yields consistent zero-shot accuracy gains on Pythia-12B downstream tasks and generalizes effectively to the Muon optimizer \citep{jordan2024muon}.
Wenjie Zhou, Bohan Wang, Hongtao Zhang +3
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, China · University of Chinese Academy of Sciences, China · Alibaba Group, China +2
Model merging aims to consolidate multiple task-specific models fine-tuned on different datasets into a unified architecture that performs cross-domain proficiency. Current data-free model merging methods often struggle to scale as they rely on simple parameter-level heuristics that ignore inter-layer dependencies and non-uniform distribution of expertise. This work proposes SA-Merging, which is built upon connectivity-based saliency formulations from structural pruning (e.g., SynFlow) and extends them to the data-free model merging setting. We define a saliency score over task vectors relative to a shared base model, and further introduce merge-aware modulation that incorporates agreement across experts to mitigate task interference. Based on this formulation, an iterative saliency-aware merging procedure progressively removes non-informative updates while preserving end-to-end connectivity. Furthermore, we extend SA-Merging to introduce rank-wise saliency decomposition for LoRAs without compromising their structural integrity. Extensive experiments on vision and language tasks demonstrate the effectiveness of our saliency-based approach, further reducing the gap between data-free and test-time adaptation methods.
Jungin Park, Jiyoung Lee, Kwanghoon Sohn
Yonsei University, Seoul, South Korea · Ewha Womans University, Seoul, South Korea.