The performance ceiling of an LLM team is constrained not only by individual model capabilities, but also by inter-member error resonance and predictive differences. Although heterogeneous teaming is often observed to be effective in practice, existing approaches lack complementarity metrics that are computable, interpretable, and optimizable, leaving team composition to rely on heuristics. We propose a heterogeneity-driven team selection framework that performs offline profiling to characterize individual capability along with two complementary signals: one captures decorrelation in error patterns to reduce co-failures, while the other measures divergence in predictive behavior to capture strategy diversity. We formulate team selection as a standardized quality--complementarity combinatorial objective and apply an efficient greedy search to select a small team from a candidate pool. Experiments across multiple benchmarks demonstrate that our framework consistently outperforms quality-only baselines under controlled candidate pools and team sizes, establishing reusable selection principles for multi-LLM systems.
Figures & tables
Figure 1 : Illustration of heterogeneity-aware team selection.
Figure 2 : Overview of C 2 -MAS framework. Given a candidate pool of m models, we (1) profile each model on development data to obtain quality estimates and choice distributions; (2) compute two heterogeneity index metrics ( HIerr and HIdist ) and stabilize them via cross-task pooling; (3) select a team of size k by optimizing a composite objective that balances quality and complementarity using multi-start greedy search; and (4) aggregate the selected team’s predictions to produce final outputs, evaluated on held-out test data.
Method
Dataset
Avg.
ARC-C
CSQA
LogiQA2
MedQA
MMLU
MMLU-Pro
OBQA
Closed-Source Models (for reference)
Gemini-2.5-Flash
92.70
81.65
60.96
77.43
81.37
58.61
92.70
77.92
GPT-4o
94.10
83.33
64.98
84.74
84.64
50.19
91.85
79.12
Single-Model Aggregation
Self-Consistency
89.14
78.65
67.32
61.70
73.41
44.76
88.58
71.94
Table 1 : Test accuracy (%) on the 7 primary benchmarks; teams of size k=3 are selected from m=11 candidates. ↑ : gain over Random- k (pp); bold : best within each aggregation block. Held-out and full results: Tables 5 and E.1 . ARC-C: ARC-Challenge; CSQA: CommonsenseQA; OBQA: OpenBookQA.
Figure 3 : Robustness and data efficiency of C 2 -MAS. (a) Test accuracy under Stacking across heterogeneity weights; (0,0) denotes Quality-Only and ⋆ marks the configuration used in our main experiments. (b–c) Test accuracy and selection stability versus profiling size, with stability measured by mean Jaccard similarity to the full-dev team.
HIerr
HIdist
Choice- Soft
PoE
DS
Stacking
Overall
✗
✗
74.79
75.00
74.26
74.72
74.69
✓
✗
75.24
75.20
74.64
75.27
75.09
✓
✓
75.45
75.24
74.61
75.52
75.21
Table 2 : Ablation study of C 2 -MAS. The first row corresponds to the Quality-Only baseline.
Figure 4 : MDS visualization of model heterogeneity.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Definition
T
Set of tasks used for cross-task pooling.
t
Task / benchmark index.
Dtdev,Dttest
Development and test splits for task t (strict dev → test).
Yt
Label / option set for task t (multiple-choice).
x,y
An instance and its gold label ( y∈Yt ).
M={1,…,m}
Index set of the m candidate agents (callable LLMs).
Appendix
Table 3: Notations and Definitions
Method
Dataset
Avg.
ARC-C
CSQA
LogiQA2
MedQA
MMLU
MMLU-Pro
OBQA
Open-Weight Individual Models
Llama-3.1-8B
80.15
75.28
54.03
61.98
67.70
37.92
81.74
65.54
Qwen3-8B
89.61
79.78
67.23
61.05
71.72
45.79
83.24
71.20
Gemma-2-9B
88.48
78.56
63.20
59.74
74.44
44.01
85.49
70.56
Ministral-3-8B
79.40
62.08
59.93
55.52
67.88
41.20
68.45
62.07
Appendix
Table E.1 : Complete per-benchmark results across all 13 benchmarks. We report the full open-weight candidate pool ( m=11 ), the single-model Self-Consistency baseline, and C 2 -MAS under four aggregation rules. (a) covers the seven primary benchmarks used in Table 1 ; (b) covers the six OOD benchmarks, where heterogeneity matrices are pooled only on the primary set and applied to OOD teams (dev → test). Avg. reports the average accuracy within each panel. Abbreviations: ARC-C: ARC-Challenge; CSQA: CommonsenseQA; OBQA: OpenBookQA.
Method
Dataset
Avg.
AQUA-RAT
C-Eval
MathQA
RACE
ReClor
StrategyQA
Open-Weight Individual Models
Llama-3.1-8B
33.99
49.53
33.24
60.77
62.45
67.32
51.22
Qwen3-8B
42.60
75.09
40.07
63.11
81.46
67.88
61.70
Gemma-2-9B
30.90
55.34
32.21
64.51
77.06
70.13
55.03
Ministral-3-8B
31.37
58.80
27.15
58.80
67.79
58.52
50.41
Appendix
Table E.1 (continued).
Method
Dataset
Avg.
AQUA-RAT
C-Eval
MathQA
RACE
ReClor
StrategyQA
Closed-Source Models (for reference)
Gemini-2.5-Flash
26.12
73.41
30.15
63.86
79.21
62.73
55.91
GPT-4o
29.96
72.47
29.49
65.07
84.36
80.06
60.24
Single-Model Aggregation
Self-Consistency
45.51
81.46
39.23
67.60
81.65
71.25
64.45
Appendix
Table 5 : OOD test accuracy (%) on the remaining 6 benchmarks. For C 2 -MAS, we pool heterogeneity matrices using only the seven primary benchmarks in Table 1 and apply them to select teams on these 6 benchmarks (dev → test). ↑ indicates gain over Random- k (pp). Bold marks the best within each aggregation block.
Benchmark
Quality-Only Team
C 2 -MAS Team
Δ Acc.
CommonsenseQA
Gemma-2 Qwen3 InternLM3
Gemma-2 Qwen3 GLM-4
↑2.43
MMLU-Pro
Nemotron Qwen3 Gemma-2
Nemotron Qwen3 Ministral-3
↑1.78
OpenBookQA
Nemotron Gemma-2 InternLM3
Nemotron Gemma-2 GLM-4
↑1.40
Appendix
Table 6 : Team composition changes made by C 2 -MAS relative to Quality-Only under Stacking aggregation. We report only benchmarks where the selected teams differ. Δ Acc. denotes the accuracy difference (pp) between C 2 -MAS and Quality-Only; arrows indicate gains or losses.
Figure 5 : Performance scaling with team size ( k ) across different aggregators. We compare the average test accuracy of C 2 -MAS (ours) against Quality-Only and Caruana baselines as the team size k varies from 2 to 5. The evaluation is performed under four distinct inference-time aggregation strategies: (a) Choice-Soft , (b) PoE , (c) DS , and (d) Stacking . C 2 -MAS demonstrates consistent improvements over baselines across all strategies, particularly in Stacking and DS , indicating that our heterogeneity-aware selection is robust to the choice of downstream combination mechanism.
Figure 6 : Validity of the optimization objective. To assess the reliability of our selection criterion, we compute the Spearman rank correlation coefficient ( ρ ) between the team scores derived from our objective function and their actual downstream test accuracy, calculated across all (311) possible team combinations for each task. The histogram illustrates the distribution of these correlations across the 13 benchmarks. With a high mean correlation of 0.751 , the results confirm that our proposed objective serves as an effective proxy for ground-truth performance, allowing C 2 -MAS to accurately rank candidate teams based solely on profiling data.
Figure 7 : Empirical stability of pooled HI under dev subsampling. We subsample 70% of each task’s dev set for 50 repetitions, recompute per-task HI matrices (Yule’s- QR and JSD), and form pooled matrices from these noisy estimates. Pooling substantially reduces the across-repetition standard deviation of pairwise entries, complementing the selection-stability analysis in Sections C.5 and C.6 .
Multi-AI collaboration, such as ensembling or debating large language models (LLMs), is a promising paradigm for aggregating information and boosting performance. A foundational step in these pipelines is to feed the responses of several proposer LLMs into a summarizer LLM, which synthesizes a better answer. However, choosing which proposers to include is non-trivial. Existing approaches primarily focus either on accuracy (picking the strongest models) or diversity (ensuring variety), and often overlook the interactions among proposers and with the summarizer. We reframe proposer selection as a combinatorial selection problem akin to feature selection, where the value of an LLM lies in its complementarity with others. However, directly applying standard feature-selection algorithms is impractical in the LLM setting due to prohibitive time complexity. Motivated by this limitation, we explore an extensive range of computationally feasible, greedy-style selection algorithms that assess complementarity using a small labeled set. Our experiments validate complementarity as a guiding principle for proposer selection and identify methods that achieve the best performance-cost trade-offs in practice.
Yichi Zhang, Kevin Lu, Yuang Zhang +3
DIMACS, Rutgers University · Department of Mathematics, Rutgers University · Department of Computer Science, George Mason University +1
Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigate whether team-based interaction improves LLM performance in What? Where? When? (ChGK), a quiz game designed to reward collective reasoning. We introduce three team strategies: Voting, Silent Team (the captain observes final answers), and Talkative Team (the captain observes both answers and rationales). To minimize data leakage, we evaluate these strategies on a dataset consisting of 572 ChGK questions released in 2025. Using six recent large-scale open models, we show that team-based strategies outperform single-model baselines, yielding gains of up to 20 percentage points in accuracy. The best team achieves 44.23% accuracy, and approaches human team performance on questions with available human statistics. Analysis of inter-model diversity reveals that disagreement strongly predicts lower accuracy, but explanatory communication substantially mitigates performance drops. We further examine captain behavior and find no evidence of self-preference bias; access to peer rationales improves captain judgments. Overall, LLM teams function primarily as answer selection and error-filtering mechanisms rather than generators of novel solutions. Our findings highlight the importance of interaction and suggest adaptive strategies as a promising direction for multi-agent systems.
Anastasia Kotelnikova, Viktor Byzov, Maria Dolzhenkova +1
Vyatka State University Kirov, Russia · European University at St. Petersburg St. Petersburg, Russia
Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on the models producing measurably different conversational behaviors when given the same input. Prior offline studies recommend drawing one model per family for behavioral diversity, because LLMs prefer outputs from their own family when rating one another in isolation. Whether the same family label predicts behavior in interactive multi-LLM systems, the setting that real deployed systems use, has not been tested. We study this with a 940,000-chain 11-checkpoint corpus and a 1.6M-chain same-base Llama factorial. On our validated headline metric, hedging, a reasoning-distilled Llama checkpoint shifts by 18% depending on which same-base partner it replies to, more than any cross-family hedging gap in the controlled subset. Qwen, closed-API, and runtime checks suggest the pattern is not isolated, while repair and challenge analyses remain exploratory because their surface-cue detectors are weaker. Overall, the results identify post-training recipe as a first-class axis for multi-LLM panel composition and show that model family alone is an incomplete proxy for conversational diversity.
Luyang Zhang, Jialu Wang, Fei Xue +1
Carnegie Mellon University · University of California, Santa Cruz · Independent Researcher