Organizations: The Hong Kong Polytechnic University (PolyU) · The Hong Kong University of Science and Technology (Guangzhou) · The Chinese University of Hong Kong · PolyU-Daya Bay Technology and Innovation Research Institute
Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.
Figures & tables
Figure 1: Overview of SMAT. (a) SMAT simulates merged parameter states during expert training through Scale, Mask, and Perturb. (b) This improves merged performance while retaining training speed close to standard fine-tuning.
Figure 2: Simulating merged updates. Scale reweights the task vector, Mask drops and rescales its coordinates, and Perturb adds uniform noise to simulate other experts’ updates.
Figure 3: Perturb alone on Llama-1B (AdamW, t=4 ). Five-merger means ± SD.
Figure 4: Efficient SMAT training. Each period of t=4 steps includes one SMAT update. Fused kernels and storage switching reduce parameter overhead.
Method \CT@row@color \CT@row@color
Expert ↑ \CT@row@color \CT@row@color
Merged Model Performance ↑ \CT@row@color \CT@row@color
Training cost ↓
\CT@row@color \CT@row@color
\CT@row@color \CT@row@color
WA
TA
TIES
DARE
DELLA
Avg. \CT@row@color \CT@row@color
Time ( × FT)
Mem. (GiB)
Llama-3.2-1B-Instruct
FT \CT@row@color \CT@row@color
54.43 \CT@row@color \CT@row@color
39.88 ↑ 0.0
42.64 ↑ 0.0
42.93 ↑ 0.0
43.50 ↑ 0.0
41.17 ↑ 0.0
42.02 ↑ 0.0 \CT@row@color \CT@row@color
1.000
20.0 ↑ 0.00%
ASAM \CT@row@color \CT@row@color
55.20 \CT@row@color \CT@row@color
38.05 ↓ 1.8
46.29 ↑ 3.6
45.48 ↑ 2.5
45.71 ↑ 2.2
44.56 ↑ 3.4
44.02 ↑ 2.0 \CT@row@color \CT@row@color
2.869
22.3 ↑ 11.5%
MergOPT \CT@row@color \CT@row@color
56.23 \CT@row@color \CT@row@color
39.67 ↓ 0.2
45.62 ↑ 3.0
45.12 ↑ 2.2
45.59 ↑ 2.1
44.54 ↑ 3.4
44.11 ↑ 2.1 \CT@row@color \CT@row@color
1.718
22.3 ↑ 11.5%
OrthoReg \CT@row@color \CT@row@color
54.44 \CT@row@color \CT@row@color
43.85 ↑ 4.0
45.77 ↑ 3.1
45.93 ↑ 3.0
45.16 ↑ 1.7
42.57 ↑ 1.4
44.66 ↑ 2.6 \CT@row@color \CT@row@color
5.693
28.2 ↑ 41.1%
Table 1: TRACE results with AdamW. Bold marks the best score for each model.
Method \CT@row@color \CT@row@color
Expert ↑ \CT@row@color \CT@row@color
Merged Model Performance ↑ \CT@row@color \CT@row@color
Training cost ↓
\CT@row@color \CT@row@color
\CT@row@color \CT@row@color
WA
TA
TIES
DARE
DELLA
Avg. \CT@row@color \CT@row@color
Time ( × FT)
Mem. (GiB)
CLIP ViT-B/32
FT \CT@row@color \CT@row@color
89.26 \CT@row@color \CT@row@color
67.22 ↑ 0.0
71.13 ↑ 0.0
73.05 ↑ 0.0
71.21 ↑ 0.0
71.75 ↑ 0.0
70.87 ↑ 0.0 \CT@row@color \CT@row@color
1.000
5.8 ↑ 0.00%
ASAM \CT@row@color \CT@row@color
90.61 \CT@row@color \CT@row@color
67.57 ↑ 0.4
72.78 ↑ 1.6
75.05 ↑ 2.0
72.78 ↑ 1.6
73.68 ↑ 1.9
72.37 ↑ 1.5 \CT@row@color \CT@row@color
1.730
6.1 ↑ 5.67%
MergOPT \CT@row@color \CT@row@color
90.05 \CT@row@color \CT@row@color
68.13 ↑ 0.9
73.15 ↑ 2.0
74.82 ↑ 1.8
72.91 ↑ 1.7
73.84 ↑ 2.1
72.57 ↑ 1.7 \CT@row@color \CT@row@color
1.146
6.1 ↑ 5.67%
OrthoReg \CT@row@color \CT@row@color
89.36 \CT@row@color \CT@row@color
70.74 ↑ 3.5
72.82 ↑ 1.7
74.37 ↑ 1.3
72.85 ↑ 1.6
70.11 ↓ 1.6
72.18 ↑ 1.3 \CT@row@color \CT@row@color
1.125
6.5 ↑ 13.8%
Table 2: CLIP results with Adam. Bold marks the best score for each model.
Method \CT@row@color \CT@row@color
Expert ↑ \CT@row@color \CT@row@color
Avg. ↑ \CT@row@color \CT@row@color
Time ↓ ( × FT)
FT \CT@row@color \CT@row@color
55.31 \CT@row@color \CT@row@color
38.25 ↑ 0.0 \CT@row@color \CT@row@color
1.00
ASAM \CT@row@color \CT@row@color
54.69 \CT@row@color \CT@row@color
40.34 ↑ 2.1 \CT@row@color \CT@row@color
1.66
MergOPT \CT@row@color \CT@row@color
54.94 \CT@row@color \CT@row@color
40.22 ↑ 2.0 \CT@row@color \CT@row@color
1.25
OrthoReg \CT@row@color \CT@row@color
50.89 \CT@row@color \CT@row@color
43.41 ↑ 5.2 \CT@row@color \CT@row@color
2.87
SMAT \CT@row@color \CT@row@color
55.07 \CT@row@color \CT@row@color
44.15 ↑ 5.9 \CT@row@color \CT@row@color
1.00
Table 3: TRACE with Muon on Llama-1B. Bold marks the best score.
Figure 5: Llama-1B expert merging with AdamW: (a) varying expert count, (b) mixing MAT and FT experts, and (c) expert scores (open) versus mean merged scores (filled). Relative scores in (a,b) are normalized to FT experts.
Figure 6: Loss slices for six tasks using the experts in Table 1 . Columns are FT, MergOPT, and SMAT; stars mark independent experts. Each task shares one NLL color scale across methods; scales differ across tasks, and contours are spaced by 0.1.
Variant
Avg. ↑
Change
Full SMAT
45.77
–
w/o Scale
44.52
−1.25
w/o Mask
44.70
−1.07
w/o Perturb
43.67
−2.10
Table 4: Component ablation on Llama-1B with AdamW. Avg. averages five mergers; Change is relative to Full SMAT.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
ASAM radius ρ
MergOPT Laplace scale bMergOPT
OrthoReg λreg
Llama-1B
2
5×10−4
0.03
Llama-8B
0.25
2.5×10−4
0.01
ViT-B/32
0.25
2×10−4
0.1
ViT-L/14
0.125
5×10−5
0.05
Appendix
Table 5: Baseline training hyperparameters for the four backbones.
Method
Expert
Merged
Avg.
WA
TA
TIES
DARE
DELLA
FT
55.31
37.06
40.46
38.53
40.92
34.26
38.25
ASAM
54.69
39.50
41.35
41.65
41.89
37.32
40.34
MergOPT
54.94
38.07
42.52
41.90
42.02
36.59
40.22
OrthoReg
50.89
42.60
44.73
43.98
45.42
40.35
43.41
SMAT
55.07
42.77
46.27
46.44
45.72
39.55
44.15
Appendix
Table 6: Muon results on Llama-3.2-1B-Instruct. Avg. is the mean over five merger methods.
Llama-1B
Llama-8B
ViT-B/32
ViT-L/14
Method
Mean
Total
Mean
Total
Mean
Total
Mean
Total
FT
0.0687
1,382.1
0.3543
7,130.1
0.1255
4,015.3
1.4494
46,381.4
ASAM
0.1971
3,965.7
0.9481
19,079.0
0.2171
6,946.1
2.8991
92,771.3
MergOPT
0.1180
2,375.2
0.5469
11,005.0
0.1438
4,601.5
1.4862
47,559.9
OrthoReg
0.3910
7,868.5
4.1047
82,599.3
0.1411
4,515.7
1.5006
48,019.5
SMAT
0.0700
1,409.0
0.3594
7,231.2
0.1268
4,057.5
1.4516
46,451.0
Appendix
Table 7: Training time for complete expert sets. Mean is seconds per optimizer step; total is summed wall time in seconds.
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, China · University of Chinese Academy of Sciences, China · Alibaba Group, China +2