Model Merging via Data-Free Covariance Estimation
Organizations: Mila & DIRO, Universit´e de Montr´eal · University of Toronto & Vector Institute
Abstract
Model merging provides a way of cheaply combining individual models to produce a model that inherits each individual's capabilities. While some merging methods can approach the performance of multitask training, they are often heuristically motivated and lack theoretical justification. A principled alternative is to pose model merging as a layer-wise optimization problem that directly minimizes interference between tasks. However, this formulation requires estimating per-layer covariance matrices from data, which may not be available when performing merging. In contrast, many of the heuristically-motivated methods do not require auxiliary data, making them practically advantageous. In this work, we revisit the interference minimization framework and show that, under certain conditions, covariance matrices can be estimated directly from difference matrices, eliminating the need for data while also reducing computational costs. We validate our approach across vision and language benchmarks on models ranging from 86M parameters to 7B parameters, outperforming previous data-free state-of-the-art merging methods
Figures & tables
| Method ( ) | HumanEval | HumanEval+ | AIME ’24 | AIME ’25 | IFEval | Avg | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| Eval ( ) | @1 | @10 | @1 | @10 | @1 | @32 | @1 | @32 | @1 | @1 |
| Zero-shot | 50.8 | 82.2 | 47.3 | 78.6 | 22.1 | 66.7 | 21.6 | 53.3 | 30.7 | 34.5 |
| Expert | 62.0 | 85.3 | 57.0 | 84.7 | 38.2 | 76.7 | 30.7 | 66.7 | 82.4 | 54.1 |
| RegMean | 59.1 | 85.7 | 54.8 | 81.1 | 30.8 | 73.3 | 27.8 | 60.0 | 74.7 | 49.4 |
| Average | 59.9 | 85.9 | 54.5 | 81.9 | 34.6 | 83.3 | 30.5 | 60.0 | 67.5 | 49.4 |
| Iso-C | 59.0 | 84.4 | 53.8 | 82.9 | 32.5 | 76.7 | 31.5 | 63.3 | 63.0 | 48.0 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Vision (OpenCLIP ViT) | Language (T5) | |
| Optimizer | AdamW | AdamW |
| Learning rate | (cosine schedule w/ -step warm-up) | , constant |
| Weight decay | ||
| Batch size | (base); (large) | |
| Gradient clipping | none | |
| Precision | float32 | bfloat16 |
| Task | OLMES task id |
|---|---|
| HumanEval | codex_humaneval::tulu |
| HumanEval+ | codex_humanevalplus::tulu |
| AIME 2024 | aime:zs_cot_r1::pass_at_32_2024_deepseek |
| AIME 2025 | aime:zs_cot_r1::pass_at_32_2025_deepseek |
| IFEval | ifeval::tulu |
| Method | ViT-B/16 | ViT-B/32 | ViT-L/14 |
|---|---|---|---|
| RegMean ( nn.MultiheadAttention ) | 83.1 | 79.4 | 87.2 |
| RegMean | 87.6 | 83.0 | 90.0 |
| Method ( ) | Data-free | NLP | Vision | |||
|---|---|---|---|---|---|---|
| Model ( ) | T5-B | T5-L | ViT-B/16 | ViT-B/32 | ViT-L/14 | |
| Zero-shot | - | 54.0 | 51.3 | 55.5 | 48.2 | 65.2 |
| Experts | - | 76.0 | 82.4 | 94.6 | 90.4 | 94.1 |
| Multitask | - | 78.7 | 82.7 | 92.1 | 89.8 | 93.7 |
| RegMean | ✗ | 74.5 | 80.8 | 87.6 | 83.0 | 90.0 |
| TA | ✗ | 71.4 | 69.7 | 76.1 | 69.7 | 84.3 |
| Method ( ) | Data-free | NLP | Vision | |||
|---|---|---|---|---|---|---|
| Model ( ) | T5-B | T5-L | ViT-B/16 | ViT-B/32 | ViT-L/14 | |
| Zero-shot | - | 54.0 | 51.3 | 55.5 | 48.2 | 65.2 |
| Experts | - | 76.5 | 79.2 | 84.1 | 80.7 | 89.0 |
| RegMean | ✗ | 75.9 | 77.7 | 80.3 | 76.5 | 86.9 |
| TA | ✗ | 66.7 | 70.9 | 75.7 | 70.5 | 83.0 |
| TA | ✓ | 54.2 | 55.0 | 66.1 | 53.5 | 77.6 |
| Method | Cars | DTD | EuroSAT | GTSRB | MNIST | RESISC45 | SUN397 | SVHN | Avg |
|---|---|---|---|---|---|---|---|---|---|
| Zero-shot | 64.7 | 44.7 | 55.8 | 43.3 | 51.8 | 66.3 | 65.5 | 52.0 | 55.5 |
| Experts | 87.5 | 98.3 | 99.7 | 99.0 | 99.8 | 97.0 | 78.8 | 97.8 | 94.7 |
| RegMean | 78.5 | 79.6 | 96.5 | 89.0 | 99.0 | 88.4 | 71.7 | 94.9 | 87.2 |
| Average | 70.1 | 54.3 | 81.1 | 60.8 | 94.4 | 75.5 | 68.6 | 72.7 | 72.2 |
| TA | 61.5 | 53.8 | 76.7 | 72.5 | 98.3 | 65.6 | 47.8 | 87.1 | 70.4 |
| TIES | 73.2 | 69.0 | 88.3 | 73.4 | 98.7 | 79.2 | 65.9 | 88.1 | 79.5 |
| Method ( ) | ViT-B/16 | ViT-B/32 | ViT-L/14 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Num. Experts ( ) | 8 | 14 | 20 | 8 | 14 | 20 | 8 | 14 | 20 |
| Zero-shot | 55.5 | 61.6 | 59.8 | 48.2 | 57.9 | 56.4 | 65.2 | 69.6 | 66.1 |
| Experts | 94.7 | 93.1 | 93.3 | 90.4 | 89.7 | 90.4 | 94.1 | 93.6 | 94.1 |
| RegMean | 87.5 | 82.2 | 78.9 | 83.0 | 78.7 | 75.4 | 90.1 | 87.6 | 85.6 |
| Average | 72.2 | 70.3 | 66.1 | 65.4 | 64.9 | 61.4 | 79.3 | 77.9 | 72.4 |
| TIES | 79.5 | 72.2 | 66.2 | 72.8 | 65.5 | 58.7 | 85.2 | 79.7 | 74.6 |
| Method | Merging FLOPs | Preprocessing FLOPs | # of Exp. Op. |
|---|---|---|---|
| Average | - | 0 | |
| Task Arith. | - | 0 | |
| RegMean | 1 | ||
| ACTMat | - | 1 | |
| Iso-C | - | 1 | |
| TSV | - | T+2 |