Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{αTransfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify αTransfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6× speedup and 70% memory reduction on vision transformers, and a 20× speedup and 85% memory reduction on large language models, while maintaining comparable performance. Our findings establish αTransfer as an efficient and generalizable approach to scaling model merging.
Figures & tables
Figure 1: Overview of α Transfer. Axes represent task vectors ( τ1,τ2 ); superscripts p , t denote proxy and target models. Contours show the performance landscape when merging with coefficients ( α1,α2 ). Optimal coefficients found on a cheap proxy model transfer directly to the large target, skipping the expensive search.
Method
The Editing Function f
Search Strategy
Granularity
Task Arithmetic ( Ilharco et al., 2023 )
Identity
Grid search
Global
TIES ( Yadav et al., 2023 )
Trim + sign election
Grid search
Global
DARE ( Yu et al., 2024 )
Random masking + rescaling
Grid search
Global
AdaMerging ( Yang et al., 2024 )
Identity
Gradient-based
Task-wise/Layer-wise
AdaMerging++ ( Yang et al., 2024 )
Trim + sign election
Gradient-based
Task-wise/Layer-wise
DivMerge ( Touayouch et al., 2026 )
Identity
Gradient-based
Task-wise
Table 1: Representative model merging methods under our unified formulation. Methods are characterized along three dimensions: editing function f , coefficient search strategy, and coefficient granularity.
Figure 3
Proxy → Target
Method
Acc (%)
Time (h)
Memory (GB)
Origin
Transfer
SA
Origin
Transfer
Speedup
Origin
Transfer
Reduction
ViT-B/32 → ViT-L/14
Task Arithmetic
75.98
75.55
72.89
11.43
1.88
6.07 ×
2.12
0.67
68.7%
TIES
75.17
71.65
64.96
12.12
2.05
5.92 ×
2.12
0.67
68.7%
DARE
76.05
75.56
62.70
13.01
2.26
5.75 ×
2.12
0.67
68.7%
AdaMerging
75.67
71.83
72.89
1.69
0.72
2.35 ×
17.84
5.27
70.4%
AdaMerging++
79.10
72.08
64.96
1.46
0.83
1.77 ×
17.84
5.27
70.4%
Table 2: Coefficient transfer on vision models. We compare direct search (Origin), α Transfer (Transfer) and simple average baseline (SA) across different proxy-target pairs. Transfer maintains accuracy comparable to the origin methods while significantly mitigating resource bottlenecks, yielding up to a 6.07 × speedup and a 70.4% memory reduction.
Proxy → Target
Method
Acc (%)
Time (h)
Memory (GB)
Origin
Transfer
SA
Origin
Transfer
Speedup
Origin
Transfer
Reduction
0.6B → 1.7B
Task Arithmetic
76.44
74.63
63.85
0.78
0.07
11.00 ×
17.21
5.96
65.4%
TIES
58.35
58.26
32.65
1.86
0.10
18.96 ×
17.21
5.96
65.4%
TIES-Pairwise
71.15
70.37
44.01
1.55
0.08
19.94 ×
10.32
3.58
65.3%
0.6B → 4B
Task Arithmetic
76.03
76.03
71.91
1.12
0.10
11.03 ×
40.22
5.96
85.2%
TIES
66.04
65.26
41.04
3.80
0.20
18.98 ×
40.22
5.96
85.2%
Table 3: Coefficient transfer on Qwen3 LLMs. We compare direct search (Origin), coefficient transfer (Transfer) and simple average baseline (SA) across different proxy-target pairs. Transfer maintains accuracy comparable to the origin methods while significantly mitigating resource bottlenecks, yielding up to a 19.97 × speedup and an 85.2% memory reduction. We include a TIES-Pairwise variant to address the poor performance of the original TIES method.
Figure 4: Budget analysis on SigLIP models. We compare the best achievable accuracy under varying computational time budgets for direct search (Origin) and α Transfer (Transfer), with simple average (SA) as a static baseline. Transfer shifts the performance curves significantly to the left, reaching comparable accuracy in a fraction of the time for most methods. The performance gap in AdaMerging++ is further investigated in Section C.1 , with full results provided in Appendix C .
Proxy → Target
Method
Acc (%)
Time Speedup
Memory Reduction
Origin
Copy
Interp
ViT-B/32 → ViT-L/14
AdaMerging
80.57
77.85
78.21
2.17 ×
70.4%
AdaMerging++
82.80
81.35
81.55
2.17 ×
70.4%
ViT-B/16 → ViT-L/14
AdaMerging
80.57
81.21
81.13
1.68 ×
61.5%
AdaMerging++
82.80
83.15
83.07
1.71 ×
61.5%
SigLIP-Base → SigLIP-Large
AdaMerging
83.57
84.25
84.21
1.53 ×
59.1%
Table 4: Layer-wise coefficient transfer. We compare direct search (Origin) against two relative-depth mapping variants: Copy and Interpolation (Interp). Remarkably, layer-wise transfer frequently outperforms direct search while reducing computational time and memory footprint.
Proxy → Target
Method
Acc (%)
Speedup
Origin
Transfer
Hybrid
Transfer
Hybrid
ViT-B/32 → ViT-L/14
Task Arithmetic
75.98
75.55
75.98
6.07 ×
1.86 ×
AdaMerging
75.67
71.83
76.00
2.35 ×
1.40 ×
DivMerge
78.77
74.73
78.51
2.14 ×
1.38 ×
ViT-B/16 → ViT-L/14
Task Arithmetic
75.98
72.89
75.98
3.52 ×
1.68 ×
AdaMerging
75.67
72.63
76.41
1.79 ×
1.28 ×
Table 5: Hybrid search results. We compare direct search (Origin), pure coefficient transfer (Transfer), and hybrid search (Hybrid) that refines transferred coefficients on the target model. Hybrid search matches or outperform direct search while still maintaining a computational speedup.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
ViT-B/32
ViT-B/16
ViT-L/14
SigLIP-B
SigLIP-L
Dataset
E
B
LR
E
B
LR
E
B
LR
E
B
LR
E
B
LR
DTD
10
32
1e-5
5
32
1e-5
10
32
1e-5
5
48
3e-5
5
48
3e-5
EuroSAT
10
32
1e-5
10
48
3e-5
10
32
1e-5
10
48
3e-5
10
48
3e-5
FER2013
10
32
1e-5
5
32
1e-5
10
32
1e-5
20
48
3e-5
20
48
3e-5
Food101
10
32
1e-5
5
32
1e-5
10
32
1e-5
20
48
3e-5
10
48
1e-5
GTSRB
10
48
3e-5
5
32
1e-5
10
32
1e-5
5
48
3e-5
5
48
3e-5
Appendix
Table 6: Selected hyperparameters for each vision model and dataset. Each cell reports the number of epochs (E), batch size (B), and learning rate (LR) chosen by the validation-accuracy criterion described in Section A.1 .
Figure 5: Budget analysis on ViT-B/32 → ViT-L/14. We compare direct search (Origin) against coefficient transfer ( α Transfer) under varying time budgets. Dashed lines indicate the simple average (SA) baseline. Top and bottom rows show evaluation- and gradient-based methods, respectively. Consistent with the SigLIP results, α Transfer shifts the performance curves to the left, reaching comparable accuracy in a fraction of the time.
Figure 6: Budget analysis on ViT-B/16 → ViT-L/14. We compare direct search (Origin) against coefficient transfer ( α Transfer) under varying time budgets. Dashed lines indicate the simple average (SA) baseline. Top and bottom rows show evaluation- and gradient-based methods, respectively. Consistent with previous observations, α Transfer shifts the performance curves to the left, reaching comparable accuracy in a fraction of the time.
Figure 7: Impact of sparsity on α Transfer in AdaMerging++. We compare exhaustive target optimization (Origin) against α Transfer across varying TIES retention ratios k∈{20,40,60,80} on the SigLIP-Base → SigLIP-Large scaling chain. The x-axis represents the wall-clock training time. The effectiveness of α Transfer is highly sensitive to the sparsity level: at extreme sparsity ( k=20 ), direct target search achieves a higher performance ceiling. However, as the retention ratio increases ( k≥60 ), α Transfer dominates the compute-performance trade-off, achieving massive speedups while converging to a higher final accuracy than direct target search.
Proxy → Target
Method
Acc (%)
Speedup
Origin
Transfer
Hybrid
Transfer
Hybrid
Qwen3-0.6B → Qwen3-1.7B
Task Arithmetic
76.44
74.63
76.44
11.00 ×
2.29 ×
TIES
58.35
58.26
58.35
18.96 ×
4.34 ×
TIES-Pairwise
71.15
70.37
71.15
19.94 ×
3.05 ×
Qwen3-0.6B → Qwen3-4B
Task Arithmetic
76.03
76.03
76.03
11.03 ×
2.30 ×
TIES
66.04
65.26
66.04
18.98 ×
4.35 ×
Appendix
Table 7: Hybrid search results on LLMs. Hybrid search recovers the accuracy of direct search while substantially reducing the computational cost.
Method
DTD
EuroSAT
FER
Food
GTSRB
RESISC
Cars
SUN
Avg
ViT-B/32 → ViT-L/14
Task Arithmetic
Or
59.63
89.07
51.67
91.50
77.53
85.40
79.59
73.48
75.98
Tr
60.53
87.11
50.35
92.62
73.34
85.59
81.10
73.77
75.55
TIES
Or
54.41
85.04
55.92
93.02
74.35
86.25
77.29
75.06
75.17
Tr
57.23
75.85
49.96
94.13
58.12
84.03
80.20
73.70
71.65
DARE
Or
60.05
88.74
51.11
92.04
76.67
85.65
80.39
73.72
76.05
Appendix
Table 8: Per-task accuracy (%) for each proxy → target pair. Or = origin (target’s best coefficients evaluated on target); Tr = transfer (proxy-best coefficients evaluated on target).
Method
DDXPlus
Banking77
Usefulness Judge
IFEval
Avg
Qwen3-0.6B → Qwen3-1.7B
Task Arithmetic
Origin
85.71
95.70
79.20
45.13
76.44
Transfer
80.78
93.50
77.60
46.64
74.63
TIES
Origin
8.56
77.10
82.40
65.34
58.35
Transfer
7.94
79.60
79.60
63.28
57.61
TIES-Pairwise
Origin
75.62
92.30
72.00
44.67
71.15
Appendix
Table 9: Per-task accuracy (%) for each LLM proxy → target pair. Origin = target’s best coefficients evaluated on target; Transfer = proxy-best coefficients evaluated on target.
Model merging combines fine-tuned checkpoints into a single multi-task model without retraining. Existing methods - such as task arithmetic, model soups, TIES, and DARE - are computationally efficient and empirically successful, but rely on heuristic design choices and lack formal optimality guarantees. We show that merging can be formulated as a convex quadratic programme over residual updates, yielding weights that minimise a squared-output calibration objective using calibration inputs and fine-tuned model outputs, and subsuming existing methods as special cases. Our framework yields a closed-form diagnostic - the fraction of residual energy captured by a chosen basis - that predicts downstream merge quality using only the calibration set. Empirically, the QP matches or outperforms existing methods in the single-layer setting, and we characterise when the optimal basis provides significant gains over the cheaper diagonal QP. We extend to multi-layer merging via a sequential layer-wise algorithm and demonstrate consistent gains across language and vision benchmarks.
Bethan Evans, Benjamin Etheridge, Stephen Roberts +1
Department of Mathematics, University of Oxford, Oxford, UK · Department of Engineering Science, University of Oxford, Oxford, UK.
Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Despite extensive prior work, most evaluations of model merging in computer vision are restricted to image classification using CLIP, where different classification datasets define different tasks. In this work, our goal is to make model merging more practical and show its relevance on challenging scenarios beyond this specific setting. In most vision scenarios, different tasks rely on trainable and usually heterogeneous decoders. Differently from previous studies with frozen decoders, where merged models can be evaluated right away, the non-trivial cost of decoder training renders hyperparameter selection based on downstream performance impractical. To address this, we introduce the task alignment proxy, and show how it can be used to speed up hyperparameter selection by orders of magnitude while retaining performance. Equipped with the task alignment proxy, we extend the applicability of model merging to multi-task vision models beyond CLIP-based classification. Project page: https://europe.naverlabs.com/task-alignment
Pau de Jorge, César Roberto de Souza, Björn Michele +5
Specialized models encode task-oriented behavior, but transferring that behavior to a general language model usually requires training, distillation, or representation alignment. We study whether such ability can instead be transferred directly at the parameter level. We apply two existing training-free heterogeneous merging methods, previously shown to transfer knowledge between general language models, to specialist-to-general transfer, projecting a specialist donor into the recipient's shape and interpolating backbone parameters without gradient updates or semantic alignment. Intersection-Merge (IM) injects a prefix-aligned donor slice matching the recipient shape, while Activate-Prune-Merge (APM) uses forward-pass activation statistics to select which donor dimensions to retain before injection. Across embedding, reranking, reward modeling, and MoE code-specialist transfer, both methods improve the general recipient, showing that simple heterogeneous merging can move capabilities across diverse specialist roles.