CASS: Contribution-Aware Structured Sparsity for Model Merging
Organizations: Pengcheng Laboratory, Shenzhen, China · City University of Hong Kong, Hong Kong, China · Southern University of Science and Technology, Shenzhen, China · Harbin Institute of Technology, Shenzhen, China · Northwestern Polytechnical University, Xi’an, China
Abstract
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, treating Transformers as unstructured ``bags of parameters'' and overlooking their inherent modularity. In this paper, we propose \textbf{C}ontribution-\textbf{A}ware \textbf{S}tructured \textbf{S}parsity (CASS), a unified framework that reduces parameter interference by identifying and preserving task-specific components. At the core of CASS is a contribution-aware structured mask that identifies task-relevant attention heads and FFN neurons. We instantiate this mask in two settings: CASS-Merging, the primary post-hoc setting where masks serve as a plug-and-play denoising filter for existing merging operators, and CASS-Tuning, an extension for scenarios with fine-tuning access where masks constrain gradients to reduce structural overlap between task vectors. Our analysis shows that task-relevant components are sparse and partially disjoint, supporting structured component-level filtering as an effective way to reduce merging interference. Extensive experiments across vision (ViT, 20 tasks) and language (RoBERTa, 8 tasks; Qwen2.5, 4 tasks) benchmarks demonstrate that CASS improves a range of representative merging baselines.
Figures & tables
| Method | ViT-B/32 | ViT-B/16 | ||
|---|---|---|---|---|
| Base | w/ CASS-M | Base | w/ CASS-M | |
| TA | 61.37 | 65.29 (+3.92) | 65.02 | 69.99 (+4.97) |
| TIES | 63.88 | 65.52 (+1.64) | 66.54 | 69.34 (+2.80) |
| DARE | 61.40 | 65.24 (+3.84) | 65.03 | 69.90 (+4.87) |
| PCB | 63.72 | 65.75 (+2.03) | 67.74 | 69.78 (+2.04) |
| TSV | 73.67 | 74.91 (+1.24) | 78.37 | 78.98 (+0.61) |
| Method | RoBERTa | Qwen2.5-0.5B | Qwen2.5-1.5B | |||
|---|---|---|---|---|---|---|
| Base | w/ CASS-M | Base | w/ CASS-M | Base | w/ CASS-M | |
| Base Model | 47.00 | – | 0.720 | – | 0.768 | – |
| Fine-Tuned | 85.91 | – | 1.000 | – | 1.000 | – |
| Post-Mask | 79.33 | – | 0.910 | – | 0.975 | – |
| TA | 66.36 | 70.12 (+3.76) | 0.743 | 0.776 (+.034) | 0.789 | 0.837 (+.048) |
| TIES | 70.05 | 71.15 (+1.10) | 0.695 | 0.773 (+.078) | 0.803 | 0.829 (+.026) |
| Method | Qwen2.5-0.5B | Qwen2.5-1.5B | ||
|---|---|---|---|---|
| Base | w/ CASS-T | Base | w/ CASS-T | |
| Base Model | 0.720 | – | 0.768 | – |
| Fine-Tuned | 1.000 | – | 1.000 | – |
| Masked FT | 0.954 | – | 0.924 | – |
| TA | 0.743 | 0.811 (+.069) | 0.789 | 0.820 (+.031) |
| TIES | 0.695 | 0.816 (+.121) | 0.803 | 0.867 (+.064) |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Capability | Fine-tuning Dataset | Evaluation Dataset | Metric |
|---|---|---|---|
| Math | DART-Math [ 47 ] | minerva_math500 [ 19 ] | math_verify accuracy |
| Instruction | TULU-3 persona IF [ 25 ] | ifeval [ 63 ] | prompt-level strict accuracy |
| Coding | Magicoder [ 52 ] | mbppplus [ 2 ] | pass@1 |
| Safety | WildGuardMix [ 15 ] | wildguardtest [ 15 ] | Micro Harm Rate ( ) |
| Hyperparameter | Value |
|---|---|
| Fine-tuning Method | LoRA |
| LoRA Rank ( ) | 32 |
| LoRA Alpha | 64 |
| LoRA Dropout | 0.1 |
| Learning Rate | |
| LR Scheduler | Cosine |
| Method | Qwen2.5-0.5B-Instruct | Qwen2.5-1.5B-Instruct | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Code | Safe | Math | Instr. | Avg. | Code | Safe | Math | Instr. | Avg. | |
| Base Model | 0.676 | 0.708 | 0.765 | 0.731 | 0.720 | 0.687 | 0.775 | 0.855 | 0.754 | 0.768 |
| Fine-Tuned | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| Post-Mask | 0.964 | 0.836 | 0.961 | 0.878 | 0.910 | 0.953 | 1.015 | 1.051 | 0.882 | 0.975 |
| TA | 0.801 | 0.687 | 0.765 | 0.718 | 0.743 | 0.735 | 0.778 | 0.942 | 0.703 | 0.789 |
| w/ CASS-M | 0.820 | 0.684 | 0.814 | 0.788 | 0.776 | 0.808 | 0.777 | 1.065 | 0.698 | 0.837 |
| Method | Qwen2.5-0.5B-Instruct | Qwen2.5-1.5B-Instruct | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Code | Safe | Math | Instr. | Avg. | Code | Safe | Math | Instr. | Avg. | |
| Base Model | 0.676 | 0.708 | 0.765 | 0.731 | 0.720 | 0.687 | 0.775 | 0.855 | 0.754 | 0.768 |
| Fine-Tuned | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| Masked FT | 0.966 | 0.988 | 0.902 | 0.962 | 0.954 | 0.885 | 1.008 | 0.978 | 0.826 | 0.924 |
| TA | 0.801 | 0.687 | 0.765 | 0.718 | 0.743 | 0.735 | 0.778 | 0.942 | 0.703 | 0.789 |
| w/ CASS-T | 0.889 | 0.705 | 0.863 | 0.788 | 0.811 | 0.798 | 0.772 | 1.007 | 0.703 | 0.820 |