Training-Free Transformer Merging via Sequential Local Operator Alignment
Organizations: CISPA Helmholtz Center for Information Security · Munich Center for Machine Learning · Technical University of Munich, School of Computation, Information and Technology · Massachusetts Institute of Technology, Department of EECS, CSAIL
Abstract
Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different activations than they do in the original model, thus, the merged and original execution paths no longer match. In this paper, we introduce Sequential Local Operator Alignment, a training-free method that merges transformers along the execution path of the partially merged model. Our method uses calibration data to estimate the local behavior of each functional component, aligns operators sequentially under the intermediate activation of the partially merged model, and subsequently factorizes the merged operators back into valid transformer parameters. We empirically show that this sequential step reduces error accumulation across layers. Furthermore, the proposed operator factorization step enables rank expansion, providing a principled mechanism for increasing multi-task capacity. We demonstrate that our approach generalizes across modalities, model scales, and varying numbers of tasks, from CLIP and RoBERTa to billion-parameter LLMs, and further extends naturally to the merging of LoRA-fine-tuned models. The results indicate improvements over strong merging baselines without requiring rank expansion, while optional expansion provides a further accuracy-inference-cost trade-off. Project link: https://akansh12.github.io/SLOA-Merge/
Figures & tables
| \begin{array}[]{c|c|c|c}\text{Component}&\text{Local operator \bm{\theta}}&\text{Behavior map \mathcal{A}({\bm{Z}};\bm{\theta})}&\text{Metric}\\ \hline\cr\text{LayerNorm}&(\bm{\gamma},\bm{\beta})&\mathcal{A}({\bm{X}};\bm{\gamma},\bm{\beta})=\nu({\bm{X}})\operatorname{diag}(\bm{\gamma})+\mathbf{1}\bm{\beta}^{\top}&\operatorname{Cov}(\nu(\chi))\\[4.2679pt] \text{Linear map}&{\bm{W}}&\mathcal{A}({\bm{X}};{\bm{W}})={\bm{X}}{\bm{W}}&{\bm{\Lambda}}\\[4.2679pt] \text{QK component}&{\bm{W}}_{Q}{\bm{W}}_{K}^{\top}/\sqrt{d_{k}}&\mathcal{A}({\bm{X}};\bm{\theta})={\bm{X}}\bm{\theta}{\bm{X}}^{\top}&{\bm{\Lambda}}\otimes{\bm{\Lambda}}\\[4.2679pt] \text{VO component}&{\bm{W}}_{V}{\bm{W}}_{O}&\mathcal{A}({\bm{A}}{\bm{X}};\bm{\theta})={\bm{A}}{\bm{X}}\bm{\theta}&{\bm{\Lambda}}\end{array} |
| Method | # Samples | Avg. Accuracy | |||
|---|---|---|---|---|---|
| 8 Tasks | 14 Tasks | 20 Tasks | |||
| Fine-tuned | – | 90.39 | 89.26 | 89.78 | |
| Data-Free | Weight Averaging | – | 66.29 | 65.37 | 61.08 |
| Task Arithmetic | – | 69.49 | 66.47 | 60.60 | |
| TIES-Merging | – | 71.70 | 67.58 | 62.76 | |
| TSV-M | – | 83.04 | 79.11 | 73.66 | |
| Method | Avg. Metric |
|---|---|
| Individual (oracle) | 0.8483 |
| Data-Free | |
| Weight Averaging | 0.6165 |
| Task Arithmetic | 0.6724 |
| TIES-Merging | 0.6448 |
| TSV-M | 0.7120 |
| Method | Benchmark | Average | |||
|---|---|---|---|---|---|
| GSM8K | HumanEval | IFEval | Avg. | ||
| Experts | Llama2-7B Pretrained | 4.32 | 0.00 | 20.89 | 8.40 |
| Llama2-7B IF-Tulu | 12.59 | 1.22 | 66.17 | 26.66 | |
| Llama2-7B Code (Magic-Coder) | 9.70 | 32.93 | 26.99 | 23.21 | |
| Llama2-7B Math (Meta-Math) | 65.50 | 0.00 | 15.53 | 27.01 | |
| Data-Free | Weight Averaging | 40.11 | 18.29 | 40.67 | 33.02 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Data Source | # Classes | Train | Test | Calib. | Calib. (%) |
| 8-Tasks | ||||||
| SUN397 | tanganke/sun397 | 397 | 76,128 | 10,875 | 256 | 0.34 |
| GTSRB | tanganke/gtsrb | 43 | 39,209 | 12,630 | 256 | 0.65 |
| Cars | tanganke/stanford_cars | 196 | 8,144 | 8,041 | 256 | 3.14 |
| EuroSAT | tanganke/eurosat | 10 | 21,600 | 5,400 | 256 | 1.19 |
| RESISC45 | tanganke/resisc45 | 45 | 25,200 | 6,300 | 256 | 1.02 |
| Task | # Labels | Primary Metric | Train | Val. | Calib. | Calib. (%) |
|---|---|---|---|---|---|---|
| CoLA | 2 | Matthews correlation | 8,551 | 1,043 | 256 | 2.99 |
| MNLI | 3 | Accuracy | 392,702 | 9,815 | 256 | 0.07 |
| MRPC | 2 | Avg. F1 & Acc. | 3,668 | 408 | 256 | 6.98 |
| QNLI | 2 | Accuracy | 104,743 | 5,463 | 256 | 0.24 |
| QQP | 2 | Avg. F1 & Acc. | 363,849 | 40,430 | 256 | 0.07 |
| RTE | 2 | Accuracy | 2,490 | 277 | 256 | 10.28 |
| Task | Train Corpus | Benchmark | Primary Metric | Train | Eval | Calib. | Calib. (%) |
|---|---|---|---|---|---|---|---|
| Math | MetaMathQA | GSM8K | Exact match (flexible extract) | 395,000 | 1,319 | 256 | 0.065 |
| Code | Magicoder-Evol-Instruct-110K | HumanEval | pass@1 | 111,183 | 164 | 256 | 0.230 |
| Instruction-following | Tulu-3 Persona-IF | IFEval | Prompt-level strict accuracy | 29,980 | 541 | 256 | 0.854 |
| Method | # Samples | Avg. Accuracy | |||
|---|---|---|---|---|---|
| 8 Tasks | 14 Tasks | 20 Tasks | |||
| Fine-tuned | – | 92.34 | 91.31 | 91.61 | |
| Data-Free | Weight Averaging | – | 72.28 | 69.72 | 64.81 |
| Task Arithmetic | – | 77.17 | 70.92 | 64.23 | |
| TIES-Merging | – | 80.21 | 71.48 | 66.31 | |
| TSV-M | – | 86.70 | 82.09 | 78.31 | |
| Method | # Samples | Norm. Acc. (%) | |
|---|---|---|---|
| Fine-tuned (LoRA, oracle) | – | 100.00 | |
| Data-Free | Simple Avg | – | 63.76 |
| TA | – | 63.70 | |
| TIES | – | 63.70 | |
| DARE-TIES | – | 63.70 | |
| KnOTS-TIES | – | 68.00 |
| Method | Data requirement | Merge time (s) | Accuracy (%) |
|---|---|---|---|
| Task Arithmetic | Data-free | 14.63 | 64.23 |
| TIES-Merging | Data-free | 39.64 | 66.31 |
| TSV-M | Data-free | 153.34 | 78.31 |
| Iso-C | Data-free | 17.09 | 75.82 |
| Iso-CTS | Data-free | 167.24 | 78.50 |
| Fisher Merging | Calibration-based ( ) | 112.70 | 66.50 |
| ViT-B/32 | ViT-B/16 | |||
|---|---|---|---|---|
| # Tasks | Accuracy (%) | Merge time (s) | Accuracy (%) | Merge time (s) |
| 8 | 85.50 0.05 | 74.2 0.5 | 88.04 0.24 | 107.50 1.76 |
| 14 | 82.61 0.04 | 103.3 3.1 | 85.16 0.09 | 136.50 0.49 |
| 20 | 81.03 0.12 | 130.6 1.3 | 83.43 0.09 | 180.49 1.88 |
| ViT-B/32 | ViT-B/16 | |||
|---|---|---|---|---|
| Accuracy (%) | Merge time (s) | Accuracy (%) | Merge time (s) | |
| 128 | 79.77 0.38 | 99.31 2.13 | 82.67 0.12 | 124.96 0.92 |
| 256 | 81.03 0.12 | 130.64 1.26 | 83.43 0.09 | 180.49 1.88 |
| 512 | 81.63 0.13 | 147.59 3.28 | 83.86 0.08 | 245.69 0.58 |
| 1024 | 81.92 0.08 | 177.78 1.89 | 84.12 0.11 | 376.18 0.57 |
| 2048 | 82.01 0.07 | 252.40 2.37 | 84.24 0.03 | 635.90 2.24 |
| Backbone | 8 tasks | 14 tasks | 20 tasks |
|---|---|---|---|
| ViT-B/32 | 2.1 | 2.5 | 2.9 |
| ViT-B/16 | 7.0 | 8.0 | 9.0 |
| Backbone | |||||
|---|---|---|---|---|---|
| ViT-B/32 | 1.4 | 2.1 | 2.4 | 3.2 | 5.5 |
| ViT-B/16 | 3.8 | 7.0 | 8.1 | 10.9 | 20.2 |
| Deployed model | QK/VO rank | Accuracy (%) | Throughput (img/s) | Latency (ms/img) | Peak memory (GB) |
|---|---|---|---|---|---|
| All baseline merges (shared original architecture) | 64 / 64 | 64.23–78.50 | 3257 | 0.307 | 0.60 |
| SLOA (original rank, ) | 64 / 64 | 82.37 | 3261.1 | 0.307 | 0.60 |
| SLOA (expanded rank, ) | 128 / 128 | 83.53 | 2711.5 | 0.369 | 0.65 |
| SLOA (expanded rank, ) | 256 / 256 | 83.74 | 1999.1 | 0.500 | 0.83 |
| Method | # Samples | Accuracy | Average | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| CoLA | SST2 | MRPC | QQP | MNLI | QNLI | RTE | Avg. | |||
| Individual | – | 0.6018 | 0.9404 | 0.8922 | 0.9141 | 0.8720 | 0.9271 | 0.7906 | 0.8483 | |
| Data-Free | Weight Averaging | – | 0.1808 | 0.8188 | 0.8116 | 0.7381 | 0.4383 | 0.7106 | 0.6173 | 0.6165 |
| Task Arithmetic | – | 0.2909 | 0.8463 | 0.8188 | 0.7869 | 0.5591 | 0.7584 | 0.6462 | 0.6724 | |
| TIES-Merging | – | 0.1951 | 0.8429 | 0.8333 | 0.8351 | 0.6501 | 0.7309 | 0.4260 | 0.6448 | |
| TSV-M | – | 0.3477 | 0.8739 | 0.8446 | 0.8286 | 0.6163 | 0.8124 | 0.6606 | 0.7120 | |
| Method | SUN397 | GTSRB | Cars | EuroSAT | RESISC45 | DTD | MNIST | SVHN | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| Fine-tuned (individual) | 74.90 | 98.91 | 78.47 | 99.07 | 95.14 | 79.79 | 99.58 | 97.27 | 90.39 |
| Simple Avg. | 65.39 | 54.92 | 62.42 | 75.67 | 70.63 | 50.53 | 86.28 | 64.49 | 66.29 |
| Task Arithmetic | 64.30 | 62.71 | 61.42 | 78.30 | 70.57 | 51.76 | 93.01 | 73.88 | 69.49 |
| TIES-Merging | 61.40 | 78.68 | 60.15 | 69.04 | 69.62 | 51.06 | 97.72 | 85.92 | 71.70 |
| TSV-M | 68.59 | 91.27 | 72.01 | 93.41 | 85.14 | 64.36 | 98.76 | 90.79 | 83.04 |
| Iso-C | 67.73 | 93.51 | 71.61 | 90.00 | 84.98 | 65.05 | 98.82 | 89.39 | 82.64 |
| Method | SUN397 | GTSRB | Cars | EuroSAT | RESISC45 | DTD | MNIST | SVHN | Flowers | PCAM | FER2013 | Pets | STL10 | CIFAR100 | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Fine-tuned (individual) | 74.90 | 98.91 | 78.47 | 99.07 | 95.14 | 79.79 | 99.58 | 97.27 | 88.62 | 87.96 | 71.62 | 92.37 | 97.55 | 88.38 | 89.26 |
| Simple Avg. | 64.73 | 45.63 | 60.40 | 66.96 | 67.06 | 47.13 | 76.57 | 50.66 | 67.25 | 65.27 | 51.62 | 84.25 | 97.20 | 70.39 | 65.37 |
| Task Arithmetic | 64.41 | 50.00 | 59.63 | 67.74 | 67.29 | 47.77 | 80.75 | 53.93 | 66.17 | 69.80 | 53.05 | 84.25 | 96.62 | 69.14 | 66.47 |
| TIES-Merging | 62.10 | 63.95 | 54.58 | 62.93 | 65.29 | 50.21 | 92.56 | 65.70 | 58.22 | 77.10 | 54.95 | 81.30 | 94.84 | 62.44 | 67.58 |
| TSV-M | 67.30 | 87.80 | 63.85 | 91.04 | 81.22 | 61.17 | 98.35 | 85.23 | 68.01 | 82.03 | 65.53 | 87.82 | 97.18 | 71.06 | 79.11 |
| Iso-C | 70.02 | 81.45 | 63.66 | 83.85 | 80.10 | 60.96 | 96.36 | 75.98 | 70.34 | 83.29 | 65.30 | 87.87 | 97.08 | 74.99 | 77.94 |
| Method | SUN397 | GTSRB | Cars | EuroSAT | RESISC45 | DTD | MNIST | SVHN | Flowers | PCAM | FER2013 | Pets | STL10 | CIFAR100 | CIFAR10 | Food101 | FMNIST | EMNIST | KMNIST | SST2 | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Fine-tuned (individual) | 74.90 | 98.91 | 78.47 | 99.07 | 95.14 | 79.79 | 99.58 | 97.27 | 88.62 | 87.96 | 71.62 | 92.37 | 97.55 | 88.38 | 97.60 | 88.40 | 94.75 | 95.62 | 98.23 | 71.28 | 89.78 |
| Simple Avg. | 64.22 | 43.14 | 59.56 | 60.85 | 64.76 | 46.44 | 71.79 | 47.22 | 66.29 | 63.95 | 50.14 | 83.92 | 96.99 | 69.80 | 92.66 | 80.38 | 71.31 | 14.98 | 11.41 | 61.78 | 61.08 |
| Task Arithmetic | 62.03 | 48.95 | 53.74 | 58.15 | 60.87 | 45.96 | 79.45 | 48.43 | 60.95 | 73.45 | 51.32 | 82.23 | 94.89 | 64.64 | 91.38 | 71.88 | 73.86 | 17.81 | 12.14 | 59.86 | 60.60 |
| TIES-Merging | 65.02 | 49.22 | 59.77 | 60.78 | 66.22 | 48.51 | 79.27 | 52.37 | 66.55 | 66.75 | 51.66 | 84.06 | 97.05 | 70.40 | 93.33 | 80.45 | 72.43 | 17.32 | 12.16 | 61.94 | 62.76 |
| TSV-M | 65.50 | 83.79 | 55.80 | 87.11 | 75.76 | 58.62 | 97.61 | 80.28 | 65.00 | 80.57 | 64.29 | 86.29 | 96.55 | 68.92 | 93.89 | 76.37 | 84.27 | 35.91 | 46.48 | 70.18 | 73.66 |
| Iso-C | 67.13 | 80.72 | 47.76 | 78.19 | 73.16 | 57.07 | 96.82 | 75.47 | 63.16 | 79.79 | 62.75 | 83.73 | 96.05 | 71.96 | 94.36 | 73.38 | 82.60 | 33.99 | 35.11 | 70.02 | 71.16 |