Unified language models are increasingly expected to combine heterogeneous capabilities, such as mathematics, code, instruction following, and controllable thinking behavior, within a single set of parameters. A common solution is sequential post-training on multiple objectives, but this entangles all objectives along one optimization trajectory and makes the final model highly sensitive to training order, data ratios, schedules, and stopping criteria. Weight-space merging offers a modular alternative, but naive merging of single-objective experts often fails: domain capabilities degrade sharply, or think/non-think modes collapse into one dominant behavior. We attribute both failures to incompatible weight-space geometry: experts trained on single objectives drift to distant regions of parameter space, placing their interpolations outside any shared low-loss basin. We propose Mixture-Trained Merging (MTM), which trains each branch on an objective-biased data mixture rather than a single objective, exposing it to cross-objective interactions and making branches compatible at merge time. MTM uses merged-model evaluations as a low-cost signal for selecting branch mixtures, avoiding expensive data-mixture ablations. The procedure is iterative: each round promotes the base model using globally selected merge coefficients and refines each branch mixture using domain-preferred coefficients under constraints that preserve other objectives. To scale beyond simplex grid search, MTM uses qNEHVI-based multi-objective Bayesian optimization. Across code, mathematics, instruction following, and think/non-think control, MTM outperforms naive merging and preserves behavioral separation where single-objective merging collapses, suggesting that effective unified models require branches trained to be mergeable.
Figures & tables
Figure 1
Figure 3: Mode-expert merging on OLMo-7B and Qwen3-4B-Base. Following Section 5 , Hard branches are trained on single response modes, while Soft branches use complementary think/non-think mixtures with a (0.7,0.3) split. Merging hard experts loses mode control under large branch imbalance, leading to outputs that follow the dominant branch (think or non-think) regardless of the prompt, whereas mixture-trained branches better preserve mode distinction.
Qwen3-4B-Base
OLMo-7B
Method
GSM8K
MATH
MBPP+
LCB
IFEval
Avg.
GSM8K
MATH
MBPP+
LCB
IFEval
Avg.
Math Expert
92.85 ± 0.22
84.40 ± 0.93
64.02 ± 1.63
20.59 ± 0.48
55.95 ± 0.39
63.56 ± 0.73
89.84 ± 0.65
79.20 ± 1.39
54.32 ± 0.55
18.58 ± 0.46
38.92 ± 1.24
56.17 ± 0.86
Code Expert
79.53 ± 6.63
53.85 ± 11.13
67.46 ± 0.49
22.03 ± 1.40
67.99 ± 0.53
58.17 ± 4.04
63.46 ± 3.72
66.20 ± 0.80
63.67 ± 0.15
16.01 ± 0.51
50.51 ± 1.29
51.97 ± 1.29
IF Expert
90.01 ± 0.47
76.40 ± 0.71
67.92 ± 0.82
18.72 ± 0.78
80.57 ± 0.50
66.72 ± 0.66
67.65 ± 2.10
52.07 ± 2.91
63.67 ± 0.81
14.61 ± 1.11
76.06 ± 1.51
54.81 ± 1.69
Linear Merge
91.91 ± 0.38
83.10 ± 0.59
67.72 ± 0.42
23.40 ± 0.29
77.26 ± 3.05
68.68 ± 0.95
83.83 ± 1.30
75.00 ± 0.53
64.55 ± 0.00
14.61 ± 0.64
61.62 ± 2.10
59.92 ± 0.91
Ties Merge
92.04 ± 0.50
83.07 ± 0.50
68.08 ± 1.46
23.35 ± 1.91
78.30 ± 1.18
68.97 ± 1.11
87.87 ± 0.39
76.20 ± 1.97
63.05 ± 1.76
14.90 ± 1.04
60.76 ± 2.35
60.56 ± 1.50
Table 1: Results for Qwen3-4B-Base and OLMo-7B under the SFT setting (3-eval pass@1 avg ± std). Bold marks the best result on the Avg. column and underline marks the second best.
Figure 4Figure 5
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Optimizer / LR
AdamW (fused), 2×10−5 , constant w/ 0.03 warmup
Weight decay / clip
0.01 / 1.0
Effective batch size
128 ( 2×16×4 GPUs)
Max sequence length
10,000
Precision
bf16 weights, tf32 matmul, grad ckpt
Distributed
DeepSpeed ZeRO-2
Plateau stop ( τ,K,β )
0.02 , 3 , 0.9 (val 512, every 20 steps)
Appendix
Table 4: SFT hyperparameters for Qwen3-4B-Base.
Optimizer / LR
AdamW (fused), 2×10−5 , constant w/ 0.03 warmup
Weight decay / clip
0.01 / 1.0
Effective batch size
64 ( 1×16×4 GPUs)
Max sequence length
10,000
Precision
bf16 weights, tf32 matmul, grad ckpt
Distributed
DeepSpeed ZeRO-2
Early stopping
disabled (fixed epoch budget)
Appendix
Table 5: SFT hyperparameters for OLMo-7B. Identical across think and nothink runs; only the input corpus and trajectory length cap differ ( 16,384 vs. 8,192 new tokens).
Table 6: Evaluation settings for each benchmark. n denotes the number of generations per prompt. All benchmarks use one greedy generation and three sampled generations with T=0.7 and top- p=0.95 across different random seeds.