cs.LGOct 5, 2026

Training-Free Transformer Merging via Sequential Local Operator Alignment

Authors: Akansh Maurya, Ya-Wei Eileen Lin, Stefanie Jegelka, Sebastian U Stich, Rotem Mulayoff

Organizations: CISPA Helmholtz Center for Information Security · Munich Center for Machine Learning · Technical University of Munich, School of Computation, Information and Technology · Massachusetts Institute of Technology, Department of EECS, CSAIL

Abstract

Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different activations than they do in the original model, thus, the merged and original execution paths no longer match. In this paper, we introduce Sequential Local Operator Alignment, a training-free method that merges transformers along the execution path of the partially merged model. Our method uses calibration data to estimate the local behavior of each functional component, aligns operators sequentially under the intermediate activation of the partially merged model, and subsequently factorizes the merged operators back into valid transformer parameters. We empirically show that this sequential step reduces error accumulation across layers. Furthermore, the proposed operator factorization step enables rank expansion, providing a principled mechanism for increasing multi-task capacity. We demonstrate that our approach generalizes across modalities, model scales, and varying numbers of tasks, from CLIP and RoBERTa to billion-parameter LLMs, and further extends naturally to the merging of LoRA-fine-tuned models. The results indicate improvements over strong merging baselines without requiring rank expansion, while optional expansion provides a further accuracy-inference-cost trade-off. Project link: https://akansh12.github.io/SLOA-Merge/

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CASS: Contribution-Aware Structured Sparsity for Model Merging

    Sep 28, 2026Yan Li, Guiping Cao, Meng Xu +5Continual Model MergingSoft-Token Representations

  2. ααTransfer: Coefficient Transfer for Efficient Model Merging

    Oct 6, 2026Shih-Cheng Huang, Zhi Rui Tam, Chieh-Yen Lin +3Transfer LearningModel Size

  3. Distilling Sequential Computation in Transformer Language Models

    Sep 23, 2026Zixuan Lan, Jessica Yang, Yanhong Li +2Transformer ArchitecturesToken Embeddings