Training-Free Transformer Merging via Sequential Local Operator Alignment
Authors: Akansh Maurya, Ya-Wei Eileen Lin, Stefanie Jegelka, Sebastian U Stich, Rotem Mulayoff
Organizations: CISPA Helmholtz Center for Information Security · Munich Center for Machine Learning · Technical University of Munich, School of Computation, Information and Technology · Massachusetts Institute of Technology, Department of EECS, CSAIL
Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different activations than they do in the original model, thus, the merged and original execution paths no longer match. In this paper, we introduce Sequential Local Operator Alignment, a training-free method that merges transformers along the execution path of the partially merged model. Our method uses calibration data to estimate the local behavior of each functional component, aligns operators sequentially under the intermediate activation of the partially merged model, and subsequently factorizes the merged operators back into valid transformer parameters. We empirically show that this sequential step reduces error accumulation across layers. Furthermore, the proposed operator factorization step enables rank expansion, providing a principled mechanism for increasing multi-task capacity. We demonstrate that our approach generalizes across modalities, model scales, and varying numbers of tasks, from CLIP and RoBERTa to billion-parameter LLMs, and further extends naturally to the merging of LoRA-fine-tuned models. The results indicate improvements over strong merging baselines without requiring rank expansion, while optional expansion provides a further accuracy-inference-cost trade-off. Project link: https://akansh12.github.io/SLOA-Merge/
Figures & tables
Figure 1 : Illustration of Sequential Local Operator Alignment . Our method aligns the effective local functional operators, merges them sequentially, and then factorizes the aligned operators. This sequential operator-level view respects coupled attention computation, reduces error accumulation across layers, and enables rank expansion to have additional multi-task capacity.
Figure 2 : Illustration of the non-invariance of independent merging in attn. operator.
Table 1 : Local functional operators and their data-dependent metrics. Here, ν(X) denotes the normalized input before the learned LayerNorm’s scale and shift, and A denotes the attention matrix
Figure 3 : Per-task comparison of SLOA ( r=128 ) with data-free and training-free baselines on ViT-B/32 with 20 vision tasks, Panels 3(a) , 3(b) and RoBERTa with 7 GLUE tasks, Panels 3(c) , 3(d) . SLOA consistently outperforms both data-free and training-free baselines.
Method
# Samples
Avg. Accuracy ↑
8 Tasks
14 Tasks
20 Tasks
Fine-tuned
–
90.39
89.26
89.78
Data-Free
Weight Averaging
–
66.29
65.37
61.08
Task Arithmetic
–
69.49
66.47
60.60
TIES-Merging
–
71.70
67.58
62.76
TSV-M
–
83.04
79.11
73.66
Table 2 : Performance comparison of model merging methods on ViT-B/32 across 8, 14, and 20 vision classification tasks. For SLOA, r denotes the retained factorization rank used for both QK and VO local operators. The setting r=64 matches the original ViT-B/32 attention-head dimension.
Method
Avg. Metric ↑
Individual (oracle)
0.8483
Data-Free
Weight Averaging
0.6165
Task Arithmetic
0.6724
TIES-Merging
0.6448
TSV-M
0.7120
Table 3 : Avg. GLUE metric for merged RoBERTa base models. Full per-task results are in Table 16 .
Method
Benchmark
Average
GSM8K
HumanEval
IFEval
Avg. ↑
Experts
Llama2-7B Pretrained
4.32
0.00
20.89
8.40
Llama2-7B IF-Tulu
12.59
1.22
66.17
26.66
Llama2-7B Code (Magic-Coder)
9.70
32.93
26.99
23.21
Llama2-7B Math (Meta-Math)
65.50
0.00
15.53
27.01
Data-Free
Weight Averaging
40.11
18.29
40.67
33.02
Table 4 : LLaMA-2-7B LoRA-expert merge results (math, code, instruction-following). Avg. is the mean score over GSM8K (math), HumanEval (code, pass@1), and IFEval (instruction-following). Bold/underline marks the best/second-best merging method per column.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Data Source
# Classes
Train
Test
Calib.
Calib. (%)
8-Tasks
SUN397
tanganke/sun397
397
76,128
10,875
256
0.34
GTSRB
tanganke/gtsrb
43
39,209
12,630
256
0.65
Cars
tanganke/stanford_cars
196
8,144
8,041
256
3.14
EuroSAT
tanganke/eurosat
10
21,600
5,400
256
1.19
RESISC45
tanganke/resisc45
45
25,200
6,300
256
1.02
Appendix
Table 5 : Data statistics of the 20 image classification datasets used for model merging, grouped by task set: the first 8 form 8-tasks the next 6 extend it to 14-tasks, and the final 6 extend it to 20-tasks.
Task
# Labels
Primary Metric
Train
Val.
Calib.
Calib. (%)
CoLA
2
Matthews correlation
8,551
1,043
256
2.99
MNLI
3
Accuracy
392,702
9,815
256
0.07
MRPC
2
Avg. F1 & Acc.
3,668
408
256
6.98
QNLI
2
Accuracy
104,743
5,463
256
0.24
QQP
2
Avg. F1 & Acc.
363,849
40,430
256
0.07
RTE
2
Accuracy
2,490
277
256
10.28
Appendix
Table 6 : Summary of GLUE tasks, Metric used, and dataset statistics
Task
Train Corpus
Benchmark
Primary Metric
Train
Eval
Calib.
Calib. (%)
Math
MetaMathQA
GSM8K
Exact match (flexible extract)
395,000
1,319
256
0.065
Code
Magicoder-Evol-Instruct-110K
HumanEval
pass@1
111,183
164
256
0.230
Instruction-following
Tulu-3 Persona-IF
IFEval
Prompt-level strict accuracy
29,980
541
256
0.854
Appendix
Table 7 : Summary of the LLM fine-tuning corpora, evaluation benchmarks, and dataset statistics. Each row is one LoRA expert (rank 32, all-linear modules) fine-tuned from LLaMA-2-7B.
Figure 4 : Prompt templates applied uniformly at training, calibration, and evaluation time for each LoRA expert. The math and instruction-following experts share the Alpaca-style template 4(a) ; the code expert uses Magicoder’s own template 4(b) .
Figure 5 : Average performance of SLOA as the number of calibration data increases for 5(a) ViT-B/32 and 5(b) ViT-B/16 on 20 vision tasks and 5(c) RoBERTa on 7 NLU tasks.
Method
# Samples
Avg. Accuracy ↑
8 Tasks
14 Tasks
20 Tasks
Fine-tuned
–
92.34
91.31
91.61
Data-Free
Weight Averaging
–
72.28
69.72
64.81
Task Arithmetic
–
77.17
70.92
64.23
TIES-Merging
–
80.21
71.48
66.31
TSV-M
–
86.70
82.09
78.31
Appendix
Table 8 : Performance comparison of model merging methods on ViT-B/16 across 8, 14, and 20 vision classification tasks. For SLOA, r denotes the retained factorization rank used for both QK and VO local operators. The setting r=64 matches the original ViT-B/16 attention-head dimension.
Method
# Samples
Norm. Acc. (%) ↑
Fine-tuned (LoRA, oracle)
–
100.00
Data-Free
Simple Avg
–
63.76
TA
–
63.70
TIES
–
63.70
DARE-TIES
–
63.70
KnOTS-TIES
–
68.00
Appendix
Table 9 : Normalized per-task accuracy of model merging methods on ViT-B/32 LoRA (KnOTS, r=16 [ Stoica et al., 2025 ] ) experts across 8 vision tasks. Normalized accuracy is the mean over tasks of (merged/fine-tuned) accuracy × 100; the fine-tuned LoRA oracle’s raw average accuracy is 84.28%.
Figure 6 : Diagnosing the source of SLOA’s gains : 6(a) error accumulation across network depth on ViT-B/16, and 6(b) merging accuracy versus the number of calibration samples on ViT-B/32, comparing SLOA at increasing rank ( r=64,128,256 ) against sequential correction in isolation (no operator composition), RegMean, and RegMean++.
Figure 16
Method
Data requirement
Merge time (s)
Accuracy (%)
Task Arithmetic
Data-free
14.63
64.23
TIES-Merging
Data-free
39.64
66.31
TSV-M
Data-free
153.34
78.31
Iso-C
Data-free
17.09
75.82
Iso-CTS
Data-free
167.24
78.50
Fisher Merging
Calibration-based ( N=256 )
112.70
66.50
Appendix
Table 10 : Wall-clock merging cost and accuracy when merging 20 task-specific ViT-B/16 checkpoints. Calibration-based methods use N=256 samples per task; data-free methods use none.
ViT-B/32
ViT-B/16
# Tasks
Accuracy (%)
Merge time (s)
Accuracy (%)
Merge time (s)
8
85.50 ± 0.05
74.2 ± 0.5
88.04 ± 0.24
107.50 ± 1.76
14
82.61 ± 0.04
103.3 ± 3.1
85.16 ± 0.09
136.50 ± 0.49
20
81.03 ± 0.12
130.6 ± 1.3
83.43 ± 0.09
180.49 ± 1.88
Appendix
Table 11 : Accuracy and merge time versus the number of merged tasks, at N=256 calibration samples per task and r=128 .
ViT-B/32
ViT-B/16
N
Accuracy (%)
Merge time (s)
Accuracy (%)
Merge time (s)
128
79.77 ± 0.38
99.31 ± 2.13
82.67 ± 0.12
124.96 ± 0.92
256
81.03 ± 0.12
130.64 ± 1.26
83.43 ± 0.09
180.49 ± 1.88
512
81.63 ± 0.13
147.59 ± 3.28
83.86 ± 0.08
245.69 ± 0.58
1024
81.92 ± 0.08
177.78 ± 1.89
84.12 ± 0.11
376.18 ± 0.57
2048
82.01 ± 0.07
252.40 ± 2.37
84.24 ± 0.03
635.90 ± 2.24
Appendix
Table 12 : Accuracy and merge time versus calibration samples per task, at 20 tasks and r=128 .
Backbone
8 tasks
14 tasks
20 tasks
ViT-B/32
2.1
2.5
2.9
ViT-B/16
7.0
8.0
9.0
Appendix
Table 13 : Peak GPU memory (GB) vs. the number of merged tasks, at N=256 calib. samples.
Backbone
N=128
N=256
N=512
N=1024
N=2048
ViT-B/32
1.4
2.1
2.4
3.2
5.5
ViT-B/16
3.8
7.0
8.1
10.9
20.2
Appendix
Table 14 : Peak GPU memory (GB) versus calib. samples per task, at 8 tasks.
Deployed model
QK/VO rank
Accuracy (%)
Throughput (img/s) ↑
Latency (ms/img) ↓
Peak memory (GB) ↓
All baseline merges (shared original architecture)
64 / 64
64.23–78.50
≈ 3257
0.307
0.60
SLOA (original rank, r=64 )
64 / 64
82.37
3261.1
0.307
0.60
SLOA (expanded rank, r=128 )
128 / 128
83.53
2711.5
0.369
0.65
SLOA (expanded rank, r=256 )
256 / 256
83.74
1999.1
0.500
0.83
Appendix
Table 15 : Inference cost by deployed rank on ViT-B/16 (20-task merge). Inference measured with batch size 64, FP16, on a single A100 GPU, averaged over 100 iterations.
Method
# Samples
Accuracy
Average
CoLA
SST2
MRPC
QQP
MNLI
QNLI
RTE
Avg. ↑
Individual
–
0.6018
0.9404
0.8922
0.9141
0.8720
0.9271
0.7906
0.8483
Data-Free
Weight Averaging
–
0.1808
0.8188
0.8116
0.7381
0.4383
0.7106
0.6173
0.6165
Task Arithmetic
–
0.2909
0.8463
0.8188
0.7869
0.5591
0.7584
0.6462
0.6724
TIES-Merging
–
0.1951
0.8429
0.8333
0.8351
0.6501
0.7309
0.4260
0.6448
TSV-M
–
0.3477
0.8739
0.8446
0.8286
0.6163
0.8124
0.6606
0.7120
Appendix
Table 16 : Multi-task performance on seven GLUE tasks with merged RoBERTa models. Training-free methods use 256 calibration samples. SLOA denotes rank-preserving operator alignment, while SLOA ( r=128 ) expands the factorization rank of the merged QK and VO operators.
Method
SUN397
GTSRB
Cars
EuroSAT
RESISC45
DTD
MNIST
SVHN
Avg.
Fine-tuned (individual)
74.90
98.91
78.47
99.07
95.14
79.79
99.58
97.27
90.39
Simple Avg.
65.39
54.92
62.42
75.67
70.63
50.53
86.28
64.49
66.29
Task Arithmetic
64.30
62.71
61.42
78.30
70.57
51.76
93.01
73.88
69.49
TIES-Merging
61.40
78.68
60.15
69.04
69.62
51.06
97.72
85.92
71.70
TSV-M
68.59
91.27
72.01
93.41
85.14
64.36
98.76
90.79
83.04
Iso-C
67.73
93.51
71.61
90.00
84.98
65.05
98.82
89.39
82.64
Appendix
Table 17 : Per-task accuracy (%) of merged ViT-B/32 models on 8 vision tasks. Training-free methods use N=256 calibration samples per task; their averages are mean ± std over three seeds. For SLOA, r is the retained rank of the QK and VO operators ( r=64 preserves the original architecture). Bold/underline marks the best/second-best merging method per column. Data-free methods are deterministic given the checkpoints, so standard deviations are reported only for calibration-based methods, computed over three random draws of the calibration set.
Method
SUN397
GTSRB
Cars
EuroSAT
RESISC45
DTD
MNIST
SVHN
Flowers
PCAM
FER2013
Pets
STL10
CIFAR100
Avg.
Fine-tuned (individual)
74.90
98.91
78.47
99.07
95.14
79.79
99.58
97.27
88.62
87.96
71.62
92.37
97.55
88.38
89.26
Simple Avg.
64.73
45.63
60.40
66.96
67.06
47.13
76.57
50.66
67.25
65.27
51.62
84.25
97.20
70.39
65.37
Task Arithmetic
64.41
50.00
59.63
67.74
67.29
47.77
80.75
53.93
66.17
69.80
53.05
84.25
96.62
69.14
66.47
TIES-Merging
62.10
63.95
54.58
62.93
65.29
50.21
92.56
65.70
58.22
77.10
54.95
81.30
94.84
62.44
67.58
TSV-M
67.30
87.80
63.85
91.04
81.22
61.17
98.35
85.23
68.01
82.03
65.53
87.82
97.18
71.06
79.11
Iso-C
70.02
81.45
63.66
83.85
80.10
60.96
96.36
75.98
70.34
83.29
65.30
87.87
97.08
74.99
77.94
Appendix
Table 18 : Per-task accuracy (%) of merged ViT-B/32 models on 14 vision tasks. Setup and notation as in Table 17 .
Method
SUN397
GTSRB
Cars
EuroSAT
RESISC45
DTD
MNIST
SVHN
Flowers
PCAM
FER2013
Pets
STL10
CIFAR100
CIFAR10
Food101
FMNIST
EMNIST
KMNIST
SST2
Avg.
Fine-tuned (individual)
74.90
98.91
78.47
99.07
95.14
79.79
99.58
97.27
88.62
87.96
71.62
92.37
97.55
88.38
97.60
88.40
94.75
95.62
98.23
71.28
89.78
Simple Avg.
64.22
43.14
59.56
60.85
64.76
46.44
71.79
47.22
66.29
63.95
50.14
83.92
96.99
69.80
92.66
80.38
71.31
14.98
11.41
61.78
61.08
Task Arithmetic
62.03
48.95
53.74
58.15
60.87
45.96
79.45
48.43
60.95
73.45
51.32
82.23
94.89
64.64
91.38
71.88
73.86
17.81
12.14
59.86
60.60
TIES-Merging
65.02
49.22
59.77
60.78
66.22
48.51
79.27
52.37
66.55
66.75
51.66
84.06
97.05
70.40
93.33
80.45
72.43
17.32
12.16
61.94
62.76
TSV-M
65.50
83.79
55.80
87.11
75.76
58.62
97.61
80.28
65.00
80.57
64.29
86.29
96.55
68.92
93.89
76.37
84.27
35.91
46.48
70.18
73.66
Iso-C
67.13
80.72
47.76
78.19
73.16
57.07
96.82
75.47
63.16
79.79
62.75
83.73
96.05
71.96
94.36
73.38
82.60
33.99
35.11
70.02
71.16
Appendix
Table 19 : Per-task accuracy (%) of merged ViT-B/32 models on 20 vision tasks. Setup and notation as in Table 17 .
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, treating Transformers as unstructured ``bags of parameters'' and overlooking their inherent modularity. In this paper, we propose \textbf{C}ontribution-\textbf{A}ware \textbf{S}tructured \textbf{S}parsity (CASS), a unified framework that reduces parameter interference by identifying and preserving task-specific components. At the core of CASS is a contribution-aware structured mask that identifies task-relevant attention heads and FFN neurons. We instantiate this mask in two settings: CASS-Merging, the primary post-hoc setting where masks serve as a plug-and-play denoising filter for existing merging operators, and CASS-Tuning, an extension for scenarios with fine-tuning access where masks constrain gradients to reduce structural overlap between task vectors. Our analysis shows that task-relevant components are sparse and partially disjoint, supporting structured component-level filtering as an effective way to reduce merging interference. Extensive experiments across vision (ViT, 20 tasks) and language (RoBERTa, 8 tasks; Qwen2.5, 4 tasks) benchmarks demonstrate that CASS improves a range of representative merging baselines.
Yan Li, Guiping Cao, Meng Xu +5
Pengcheng Laboratory, Shenzhen, China · City University of Hong Kong, Hong Kong, China · Southern University of Science and Technology, Shenzhen, China +2
Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{αTransfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify αTransfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6× speedup and 70% memory reduction on vision transformers, and a 20× speedup and 85% memory reduction on large language models, while maintaining comparable performance. Our findings establish αTransfer as an efficient and generalizable approach to scaling model merging.
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or frequently occur as stable units, suggesting that their representations may be compressible. We introduce a method for distilling sequential computation by replacing spans of input tokens with collapsed representations, computed on the fly by a lightweight merge module. This module generates a single surrogate embedding from a sequence of static token embeddings that captures the functional role of the multiple tokens, allowing pretrained models to operate on compressed inputs without architectural changes or re-training. We apply this approach during inference to compress both prompts and intermediate decoding steps, using a rollback mechanism to substitute stored multi-token KV cache entries with their single-step surrogates. Experiments across diverse models show that the merge module can be used to reduce effective sequence length by up to 40% with minimal accuracy degradation across language modeling evaluations and downstream tasks, including question answering, summarization, commonsense reasoning, and long-form mathematical reasoning. Additional lightweight adaptation of the merge module further improves the accuracy-compression trade-off in selected settings. These results demonstrate that sequential token computation in Transformers can be effectively approximated through condensed surrogate representations that approximate the original behavior without model updating.
Zixuan Lan, Jessica Yang, Yanhong Li +2
The University of Chicago · Toyota Technological Institute at Chicago · Independent Researcher +1