The performance ceiling of an LLM team is constrained not only by individual model capabilities, but also by inter-member error resonance and predictive differences. Although heterogeneous teaming is often observed to be effective in practice, existing approaches lack complementarity metrics that are computable, interpretable, and optimizable, leaving team composition to rely on heuristics. We propose a heterogeneity-driven team selection framework that performs offline profiling to characterize individual capability along with two complementary signals: one captures decorrelation in error patterns to reduce co-failures, while the other measures divergence in predictive behavior to capture strategy diversity. We formulate team selection as a standardized quality--complementarity combinatorial objective and apply an efficient greedy search to select a small team from a candidate pool. Experiments across multiple benchmarks demonstrate that our framework consistently outperforms quality-only baselines under controlled candidate pools and team sizes, establishing reusable selection principles for multi-LLM systems.
Figures & tables
Figure 1 : Illustration of heterogeneity-aware team selection.
Figure 2 : Overview of C 2 -MAS framework. Given a candidate pool of m models, we (1) profile each model on development data to obtain quality estimates and choice distributions; (2) compute two heterogeneity index metrics ( HIerr and HIdist ) and stabilize them via cross-task pooling; (3) select a team of size k by optimizing a composite objective that balances quality and complementarity using multi-start greedy search; and (4) aggregate the selected team’s predictions to produce final outputs, evaluated on held-out test data.
Method
Dataset
Avg.
ARC-C
CSQA
LogiQA2
MedQA
MMLU
MMLU-Pro
OBQA
Closed-Source Models (for reference)
Gemini-2.5-Flash
92.70
81.65
60.96
77.43
81.37
58.61
92.70
77.92
GPT-4o
94.10
83.33
64.98
84.74
84.64
50.19
91.85
79.12
Single-Model Aggregation
Self-Consistency
89.14
78.65
67.32
61.70
73.41
44.76
88.58
71.94
Table 1 : Test accuracy (%) on the 7 primary benchmarks; teams of size k=3 are selected from m=11 candidates. ↑ : gain over Random- k (pp); bold : best within each aggregation block. Held-out and full results: Tables 5 and E.1 . ARC-C: ARC-Challenge; CSQA: CommonsenseQA; OBQA: OpenBookQA.
Figure 3 : Robustness and data efficiency of C 2 -MAS. (a) Test accuracy under Stacking across heterogeneity weights; (0,0) denotes Quality-Only and ⋆ marks the configuration used in our main experiments. (b–c) Test accuracy and selection stability versus profiling size, with stability measured by mean Jaccard similarity to the full-dev team.
HIerr
HIdist
Choice- Soft
PoE
DS
Stacking
Overall
✗
✗
74.79
75.00
74.26
74.72
74.69
✓
✗
75.24
75.20
74.64
75.27
75.09
✓
✓
75.45
75.24
74.61
75.52
75.21
Table 2 : Ablation study of C 2 -MAS. The first row corresponds to the Quality-Only baseline.
Figure 4 : MDS visualization of model heterogeneity.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Definition
T
Set of tasks used for cross-task pooling.
t
Task / benchmark index.
Dtdev,Dttest
Development and test splits for task t (strict dev → test).
Yt
Label / option set for task t (multiple-choice).
x,y
An instance and its gold label ( y∈Yt ).
M={1,…,m}
Index set of the m candidate agents (callable LLMs).
Appendix
Table 3: Notations and Definitions
Method
Dataset
Avg.
ARC-C
CSQA
LogiQA2
MedQA
MMLU
MMLU-Pro
OBQA
Open-Weight Individual Models
Llama-3.1-8B
80.15
75.28
54.03
61.98
67.70
37.92
81.74
65.54
Qwen3-8B
89.61
79.78
67.23
61.05
71.72
45.79
83.24
71.20
Gemma-2-9B
88.48
78.56
63.20
59.74
74.44
44.01
85.49
70.56
Ministral-3-8B
79.40
62.08
59.93
55.52
67.88
41.20
68.45
62.07
Appendix
Table E.1 : Complete per-benchmark results across all 13 benchmarks. We report the full open-weight candidate pool ( m=11 ), the single-model Self-Consistency baseline, and C 2 -MAS under four aggregation rules. (a) covers the seven primary benchmarks used in Table 1 ; (b) covers the six OOD benchmarks, where heterogeneity matrices are pooled only on the primary set and applied to OOD teams (dev → test). Avg. reports the average accuracy within each panel. Abbreviations: ARC-C: ARC-Challenge; CSQA: CommonsenseQA; OBQA: OpenBookQA.
Method
Dataset
Avg.
AQUA-RAT
C-Eval
MathQA
RACE
ReClor
StrategyQA
Open-Weight Individual Models
Llama-3.1-8B
33.99
49.53
33.24
60.77
62.45
67.32
51.22
Qwen3-8B
42.60
75.09
40.07
63.11
81.46
67.88
61.70
Gemma-2-9B
30.90
55.34
32.21
64.51
77.06
70.13
55.03
Ministral-3-8B
31.37
58.80
27.15
58.80
67.79
58.52
50.41
Appendix
Table E.1 (continued).
Method
Dataset
Avg.
AQUA-RAT
C-Eval
MathQA
RACE
ReClor
StrategyQA
Closed-Source Models (for reference)
Gemini-2.5-Flash
26.12
73.41
30.15
63.86
79.21
62.73
55.91
GPT-4o
29.96
72.47
29.49
65.07
84.36
80.06
60.24
Single-Model Aggregation
Self-Consistency
45.51
81.46
39.23
67.60
81.65
71.25
64.45
Appendix
Table 5 : OOD test accuracy (%) on the remaining 6 benchmarks. For C 2 -MAS, we pool heterogeneity matrices using only the seven primary benchmarks in Table 1 and apply them to select teams on these 6 benchmarks (dev → test). ↑ indicates gain over Random- k (pp). Bold marks the best within each aggregation block.
Benchmark
Quality-Only Team
C 2 -MAS Team
Δ Acc.
CommonsenseQA
Gemma-2 Qwen3 InternLM3
Gemma-2 Qwen3 GLM-4
↑2.43
MMLU-Pro
Nemotron Qwen3 Gemma-2
Nemotron Qwen3 Ministral-3
↑1.78
OpenBookQA
Nemotron Gemma-2 InternLM3
Nemotron Gemma-2 GLM-4
↑1.40
Appendix
Table 6 : Team composition changes made by C 2 -MAS relative to Quality-Only under Stacking aggregation. We report only benchmarks where the selected teams differ. Δ Acc. denotes the accuracy difference (pp) between C 2 -MAS and Quality-Only; arrows indicate gains or losses.
Figure 5 : Performance scaling with team size ( k ) across different aggregators. We compare the average test accuracy of C 2 -MAS (ours) against Quality-Only and Caruana baselines as the team size k varies from 2 to 5. The evaluation is performed under four distinct inference-time aggregation strategies: (a) Choice-Soft , (b) PoE , (c) DS , and (d) Stacking . C 2 -MAS demonstrates consistent improvements over baselines across all strategies, particularly in Stacking and DS , indicating that our heterogeneity-aware selection is robust to the choice of downstream combination mechanism.
Figure 6 : Validity of the optimization objective. To assess the reliability of our selection criterion, we compute the Spearman rank correlation coefficient ( ρ ) between the team scores derived from our objective function and their actual downstream test accuracy, calculated across all (311) possible team combinations for each task. The histogram illustrates the distribution of these correlations across the 13 benchmarks. With a high mean correlation of 0.751 , the results confirm that our proposed objective serves as an effective proxy for ground-truth performance, allowing C 2 -MAS to accurately rank candidate teams based solely on profiling data.
Figure 7 : Empirical stability of pooled HI under dev subsampling. We subsample 70% of each task’s dev set for 50 repetitions, recompute per-task HI matrices (Yule’s- QR and JSD), and form pooled matrices from these noisy estimates. Pooling substantially reduces the across-repetition standard deviation of pairwise entries, complementing the selection-stability analysis in Sections C.5 and C.6 .