Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{αTransfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify αTransfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6× speedup and 70% memory reduction on vision transformers, and a 20× speedup and 85% memory reduction on large language models, while maintaining comparable performance. Our findings establish αTransfer as an efficient and generalizable approach to scaling model merging.
Figures & tables
Figure 1: Overview of α Transfer. Axes represent task vectors ( τ1,τ2 ); superscripts p , t denote proxy and target models. Contours show the performance landscape when merging with coefficients ( α1,α2 ). Optimal coefficients found on a cheap proxy model transfer directly to the large target, skipping the expensive search.
Method
The Editing Function f
Search Strategy
Granularity
Task Arithmetic ( Ilharco et al., 2023 )
Identity
Grid search
Global
TIES ( Yadav et al., 2023 )
Trim + sign election
Grid search
Global
DARE ( Yu et al., 2024 )
Random masking + rescaling
Grid search
Global
AdaMerging ( Yang et al., 2024 )
Identity
Gradient-based
Task-wise/Layer-wise
AdaMerging++ ( Yang et al., 2024 )
Trim + sign election
Gradient-based
Task-wise/Layer-wise
DivMerge ( Touayouch et al., 2026 )
Identity
Gradient-based
Task-wise
Table 1: Representative model merging methods under our unified formulation. Methods are characterized along three dimensions: editing function f , coefficient search strategy, and coefficient granularity.
Figure 3
Proxy → Target
Method
Acc (%)
Time (h)
Memory (GB)
Origin
Transfer
SA
Origin
Transfer
Speedup
Origin
Transfer
Reduction
ViT-B/32 → ViT-L/14
Task Arithmetic
75.98
75.55
72.89
11.43
1.88
6.07 ×
2.12
0.67
68.7%
TIES
75.17
71.65
64.96
12.12
2.05
5.92 ×
2.12
0.67
68.7%
DARE
76.05
75.56
62.70
13.01
2.26
5.75 ×
2.12
0.67
68.7%
AdaMerging
75.67
71.83
72.89
1.69
0.72
2.35 ×
17.84
5.27
70.4%
AdaMerging++
79.10
72.08
64.96
1.46
0.83
1.77 ×
17.84
5.27
70.4%
Table 2: Coefficient transfer on vision models. We compare direct search (Origin), α Transfer (Transfer) and simple average baseline (SA) across different proxy-target pairs. Transfer maintains accuracy comparable to the origin methods while significantly mitigating resource bottlenecks, yielding up to a 6.07 × speedup and a 70.4% memory reduction.
Proxy → Target
Method
Acc (%)
Time (h)
Memory (GB)
Origin
Transfer
SA
Origin
Transfer
Speedup
Origin
Transfer
Reduction
0.6B → 1.7B
Task Arithmetic
76.44
74.63
63.85
0.78
0.07
11.00 ×
17.21
5.96
65.4%
TIES
58.35
58.26
32.65
1.86
0.10
18.96 ×
17.21
5.96
65.4%
TIES-Pairwise
71.15
70.37
44.01
1.55
0.08
19.94 ×
10.32
3.58
65.3%
0.6B → 4B
Task Arithmetic
76.03
76.03
71.91
1.12
0.10
11.03 ×
40.22
5.96
85.2%
TIES
66.04
65.26
41.04
3.80
0.20
18.98 ×
40.22
5.96
85.2%
Table 3: Coefficient transfer on Qwen3 LLMs. We compare direct search (Origin), coefficient transfer (Transfer) and simple average baseline (SA) across different proxy-target pairs. Transfer maintains accuracy comparable to the origin methods while significantly mitigating resource bottlenecks, yielding up to a 19.97 × speedup and an 85.2% memory reduction. We include a TIES-Pairwise variant to address the poor performance of the original TIES method.
Figure 4: Budget analysis on SigLIP models. We compare the best achievable accuracy under varying computational time budgets for direct search (Origin) and α Transfer (Transfer), with simple average (SA) as a static baseline. Transfer shifts the performance curves significantly to the left, reaching comparable accuracy in a fraction of the time for most methods. The performance gap in AdaMerging++ is further investigated in Section C.1 , with full results provided in Appendix C .
Proxy → Target
Method
Acc (%)
Time Speedup
Memory Reduction
Origin
Copy
Interp
ViT-B/32 → ViT-L/14
AdaMerging
80.57
77.85
78.21
2.17 ×
70.4%
AdaMerging++
82.80
81.35
81.55
2.17 ×
70.4%
ViT-B/16 → ViT-L/14
AdaMerging
80.57
81.21
81.13
1.68 ×
61.5%
AdaMerging++
82.80
83.15
83.07
1.71 ×
61.5%
SigLIP-Base → SigLIP-Large
AdaMerging
83.57
84.25
84.21
1.53 ×
59.1%
Table 4: Layer-wise coefficient transfer. We compare direct search (Origin) against two relative-depth mapping variants: Copy and Interpolation (Interp). Remarkably, layer-wise transfer frequently outperforms direct search while reducing computational time and memory footprint.
Proxy → Target
Method
Acc (%)
Speedup
Origin
Transfer
Hybrid
Transfer
Hybrid
ViT-B/32 → ViT-L/14
Task Arithmetic
75.98
75.55
75.98
6.07 ×
1.86 ×
AdaMerging
75.67
71.83
76.00
2.35 ×
1.40 ×
DivMerge
78.77
74.73
78.51
2.14 ×
1.38 ×
ViT-B/16 → ViT-L/14
Task Arithmetic
75.98
72.89
75.98
3.52 ×
1.68 ×
AdaMerging
75.67
72.63
76.41
1.79 ×
1.28 ×
Table 5: Hybrid search results. We compare direct search (Origin), pure coefficient transfer (Transfer), and hybrid search (Hybrid) that refines transferred coefficients on the target model. Hybrid search matches or outperform direct search while still maintaining a computational speedup.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
ViT-B/32
ViT-B/16
ViT-L/14
SigLIP-B
SigLIP-L
Dataset
E
B
LR
E
B
LR
E
B
LR
E
B
LR
E
B
LR
DTD
10
32
1e-5
5
32
1e-5
10
32
1e-5
5
48
3e-5
5
48
3e-5
EuroSAT
10
32
1e-5
10
48
3e-5
10
32
1e-5
10
48
3e-5
10
48
3e-5
FER2013
10
32
1e-5
5
32
1e-5
10
32
1e-5
20
48
3e-5
20
48
3e-5
Food101
10
32
1e-5
5
32
1e-5
10
32
1e-5
20
48
3e-5
10
48
1e-5
GTSRB
10
48
3e-5
5
32
1e-5
10
32
1e-5
5
48
3e-5
5
48
3e-5
Appendix
Table 6: Selected hyperparameters for each vision model and dataset. Each cell reports the number of epochs (E), batch size (B), and learning rate (LR) chosen by the validation-accuracy criterion described in Section A.1 .
Figure 5: Budget analysis on ViT-B/32 → ViT-L/14. We compare direct search (Origin) against coefficient transfer ( α Transfer) under varying time budgets. Dashed lines indicate the simple average (SA) baseline. Top and bottom rows show evaluation- and gradient-based methods, respectively. Consistent with the SigLIP results, α Transfer shifts the performance curves to the left, reaching comparable accuracy in a fraction of the time.
Figure 6: Budget analysis on ViT-B/16 → ViT-L/14. We compare direct search (Origin) against coefficient transfer ( α Transfer) under varying time budgets. Dashed lines indicate the simple average (SA) baseline. Top and bottom rows show evaluation- and gradient-based methods, respectively. Consistent with previous observations, α Transfer shifts the performance curves to the left, reaching comparable accuracy in a fraction of the time.
Figure 7: Impact of sparsity on α Transfer in AdaMerging++. We compare exhaustive target optimization (Origin) against α Transfer across varying TIES retention ratios k∈{20,40,60,80} on the SigLIP-Base → SigLIP-Large scaling chain. The x-axis represents the wall-clock training time. The effectiveness of α Transfer is highly sensitive to the sparsity level: at extreme sparsity ( k=20 ), direct target search achieves a higher performance ceiling. However, as the retention ratio increases ( k≥60 ), α Transfer dominates the compute-performance trade-off, achieving massive speedups while converging to a higher final accuracy than direct target search.
Proxy → Target
Method
Acc (%)
Speedup
Origin
Transfer
Hybrid
Transfer
Hybrid
Qwen3-0.6B → Qwen3-1.7B
Task Arithmetic
76.44
74.63
76.44
11.00 ×
2.29 ×
TIES
58.35
58.26
58.35
18.96 ×
4.34 ×
TIES-Pairwise
71.15
70.37
71.15
19.94 ×
3.05 ×
Qwen3-0.6B → Qwen3-4B
Task Arithmetic
76.03
76.03
76.03
11.03 ×
2.30 ×
TIES
66.04
65.26
66.04
18.98 ×
4.35 ×
Appendix
Table 7: Hybrid search results on LLMs. Hybrid search recovers the accuracy of direct search while substantially reducing the computational cost.
Method
DTD
EuroSAT
FER
Food
GTSRB
RESISC
Cars
SUN
Avg
ViT-B/32 → ViT-L/14
Task Arithmetic
Or
59.63
89.07
51.67
91.50
77.53
85.40
79.59
73.48
75.98
Tr
60.53
87.11
50.35
92.62
73.34
85.59
81.10
73.77
75.55
TIES
Or
54.41
85.04
55.92
93.02
74.35
86.25
77.29
75.06
75.17
Tr
57.23
75.85
49.96
94.13
58.12
84.03
80.20
73.70
71.65
DARE
Or
60.05
88.74
51.11
92.04
76.67
85.65
80.39
73.72
76.05
Appendix
Table 8: Per-task accuracy (%) for each proxy → target pair. Or = origin (target’s best coefficients evaluated on target); Tr = transfer (proxy-best coefficients evaluated on target).
Method
DDXPlus
Banking77
Usefulness Judge
IFEval
Avg
Qwen3-0.6B → Qwen3-1.7B
Task Arithmetic
Origin
85.71
95.70
79.20
45.13
76.44
Transfer
80.78
93.50
77.60
46.64
74.63
TIES
Origin
8.56
77.10
82.40
65.34
58.35
Transfer
7.94
79.60
79.60
63.28
57.61
TIES-Pairwise
Origin
75.62
92.30
72.00
44.67
71.15
Appendix
Table 9: Per-task accuracy (%) for each LLM proxy → target pair. Origin = target’s best coefficients evaluated on target; Transfer = proxy-best coefficients evaluated on target.