A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on which models' responses once differed in quality may later be answered equally well, while a previously difficult instruction may begin to provide a useful learning signal. We propose Stackelberg Alignment, a game-theory-inspired leader-follower framework that turns instruction selection into an adaptive curriculum. An EXP3 bandit acts as the leader, allocating a fixed sampling budget across instructions and updating its sampling distribution using a reward that combines instruction difficulty and response discriminability. The language models act as followers: they respond to the selected instructions, evaluate one another's responses, and learn from the resulting preference signals through DPO or GRPO. The framework uses Elo-style reputation-weighted peer judgment and reputation-based opponent matching to support reliable and competitive model interactions. Experiments across three heterogeneous model pools and 12 benchmarks spanning scientific discovery, reasoning, code, instruction following, and knowledge show that Stackelberg Alignment achieves the highest macro-average across three diverse model pools, outperforming the strongest training-time baseline by up to 7.4% and the best static inference baseline by 12-25%. Analysis confirms that the adaptive leader concentrates duels on the most informative instructions, and ablations show that both reputation-weighted judgment and reputation-based matching improve the effectiveness of multi-LLM evolution.
Figures & tables
Figure 1: Overview of Stackelberg Alignment for one iteration. The EXP3 leader maintains learned weights wt over the instruction pool and selects instruction xk non-uniformly. Two follower models ( M3 , M4 in the figure) are matched by reputation and duel on xk , generating responses y3 and y4 . Remaining models act as peer judges, scoring both responses; scores are aggregated weighted by judge reputation, determining the winner. The Update step performs three actions: (1) reputation scores are adjusted, (2) leader weights are updated via EXP3 based on duel informativeness, and (3) the winning preference pair is added to the dataset P for DPO training, or the selected instruction and online judge scores are used directly as the reward for GRPO training. Dashed arrows show the resulting feedback loop: the updated leader weights reshape the instruction distribution, and the updated reputations and retrained models feed back into the follower pool for the next iteration.
Scientific
Reasoning
Code
Instruction
Knowledge
Method
BixBench
LabBench
SMDD
AssayBench
GPQA-Dia
MATH
HumanEval
MBPP
AlpacaEval
IFEval
TruthfulQA
CulturalBench
Avg
Pool 1: Specialized Expert LLMs
Sparta Alignment
0.301
0.308
0.168
0.017
0.313
0.820
0.719
0.634
7.528
0.707
0.686
0.656
0.524
Majority Vote
0.184
0.272
–
–
0.263
0.778
–
–
–
–
0.616
0.670
0.464
AggLM
0.291
0.317
0.007
0.029
0.273
0.817
0.649
0.565
− 2.647
0.469
0.605
0.591
0.407
Trained Router
0.233
0.263
0.011
0.015
0.313
0.735
0.623
0.511
1.718
0.631
0.605
0.633
0.428
Table 1: Performance across three model pools and 12 datasets grouped by domain. Bold : best per column within each pool; underline : second best; –: not applicable. Shaded rows ( ) are Stackelberg Alignment variants. The Avg column is the macro-average over all available datasets, with AlpacaEval min-max normalized to [0,1] .
Figure 2: Per-model eval scores at best iteration (Pool 2, Stackelberg GRPO). Row-normalized; green = best, red = worst per task.
Figure 3: Reputation score trajectories over training iterations for TruthfulQA and HumanEval (Pool 2, Stackelberg GRPO). Reputation scores diverge from the same initialization and stabilize by iteration 4–5. The leading model differs across tasks, confirming genuine task-specific competitive structure and collaborative learning landscape.
Figure 4: Normalized opponent weight assigned by the Stackelberg leader over iterations (Pool 2, Stackelberg GRPO). Dashed line = uniform selection baseline. The leader shields the emerging dominant model from being excessively used as an opponent and adapts its strategy per task.
Figure 5: Leader instruction weights over training iterations (the MATH dataset, model pool 1). Left: heatmap of normalized weight deviation from uniform for the top-32 most dynamic instructions. Right: trajectory of top-5 and bottom-5 instructions by final-iteration weight; dashed line = uniform. The leader converges to an implicit difficulty curriculum by iteration 5.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Stackelberg (GRPO)
Stackelberg (DPO)
Task
Score
95% CI
Score
95% CI
Pool 1: Specialized Expert LLMs
BixBench
0.350∗
[0.262, 0.447]
0.233
[0.155, 0.320]
LabBench
0.322∗
[0.284, 0.360]
0.311
[0.273, 0.348]
GPQA-Dia
0.354
[0.263, 0.444]
0.364
[0.273, 0.455]
MATH
0.877∗∗
[0.856, 0.898]
0.857∗∗
[0.834, 0.879]
Appendix
Table 2: Per-task scores and 95% confidence intervals for Stackelberg (GRPO) and Stackelberg (DPO). CIs computed via 10,000 Monte Carlo draws (Bernoulli for binary tasks; Gaussian for AlpacaEval). ∗ : point estimate exceeds every baseline. ∗∗ : CI lower bound exceeds every baseline. SMDD and AssayBench omitted ( CI≈±0.55 ).
Checkpoint source
Eval Task
Best Untrained
Same-Task
Math-Trained
HumanEval-Trained
MATH
0.863
0.865
–
0.866
HumanEval
0.790
0.790 (no gain)
0.974 (+18pp)
–
MBPP
0.639
0.632 (no gain)
0.963 (+33pp)
0.961 (+32pp)
Appendix
Table 3: Cross-task transfer in Pool 2. Each trained column shows the best single-model checkpoint from competitive pool training on that source task, evaluated on the target. Baseline: best untrained model in the pool per task. “No gain” indicates same-task training matched or fell below the untrained baseline.
Figure 6: Pairwise win-rate matrix for the MATH dataset (Pool 2, Stackelberg GRPO). No model dominates; win rates span 0.1–0.6; ∼ 41% of duels are draws given the objective nature of the task.
Figure 10
Figure 9: Response properties over training iterations (MATH, Pool 2, 8 models). Left: type-token ratio (lexical diversity). Right: mean response length. Bold line = pool average. No model collapses toward repetitive or length-degenerate outputs; pool diversity is maintained throughout.
Method
AlpacaEval
GPQA
MBPP
TruthfulQA
Gemini Leader + DPO
1.65
0.263
0.581
0.682
Prob. Leader + DPO
8.081
0.364
0.643
0.682
Prob. Leader + GRPO
7.819
0.353
0.653
0.684
Appendix
Table 4: Leader strategy comparison on Pool 1 across four tasks. Best score per task is bolded . AlpacaEval is scored by a reward model (higher = better); all other tasks report accuracy.
ID
HuggingFace identifier
Specialty
M0
chtmp223/Qwen2.5-7B-CLIPPER
Reasoning/alignment
M1
chengq9/ToolRL-Qwen2.5-3B
Tool use / RL
M2
AgentFlow/agentflowplanner-7b
Agent planning
M3
nanami/ladder-last16L-llama3.1-8b-instruct-sft4k
Instruction following
M4
viswavi/qwen2.5-rlcf
RL from feedback
M5
milli19/promptmii-llama3.1-8b-instruct
Prompt optimization
Appendix
Table 5: Pool 2 model configurations (LLMs from Diverse Academic Research).
Benchmark
Dev
Test
BixBench
102
103
LabBench
578
578
SMDD
135
137
AssayBench
218
334
GPQA-Diamond
99
99
MATH
956
956
Appendix
Table 6: Dataset sizes. Dev: full development set; 80% forms the instruction pool for training and 20% is held out for validation. Test: held-out evaluation set.
Method
Forward passes
Training steps
Stackelberg Alignment
O(D⋅m)
O(m)
Sparta Alignment
O(D⋅m)
O(m)
Multiagent FT
O(N⋅m)
O(m)
AggLM
O(N⋅m)
O(1)
Trained Router
O(N⋅m)
O(1)
Mixture of Agents (MoA)
O(N⋅m)
0
Appendix
Table 7: Per-iteration theoretical complexity. m : pool size; D : duels per iteration; N : dataset size; K : instruction pool size. “Inference only” methods have S=0 .
No single Large Language Model (LLM) is uniformly reliable across queries, motivating multi-model inference systems that either route among models or combine their outputs. However, routing stops after selecting an initial model, while dense collaboration invokes peers on every query. We show that collaboration is non-monotonic: peers can recover failures that no model solves alone, but can also corrupt initially correct answers. We introduce COMED (Controlled Model Escalation for Multi-LLM Deliberation), a post-anchor controller for selective cross-model collaboration. COMED uses anchor self-consistency, router margin, and a lightweight peer probe to accept confident answers, verify ambiguous cases, and escalate only when collaboration is likely beneficial. We formalize this trade-off with a rescue-harm decomposition showing that selective collaboration improves when rescued errors outweigh collaboration-induced harms. Across medical, scientific, and general reasoning benchmarks, COMED improves fixed and routed anchors in all 16 open-weight settings, with gains up to +10.7 percentage points on MedQA while invoking fewer models and using fewer decoded tokens than dense collaboration. On HLE with frontier models, COMED improves GPT-5.5 from 23.1% to 28.1%, outperforming dense collaboration and achieving the best results.
Users increasingly face the challenge of selecting an appropriate LLM for a given task from a rapidly growing pool of LLMs, each with distinct but often opaque latent properties. Compounding this challenge, users may lack the vocabulary or awareness to explicitly articulate the characteristics they value in an LLM's responses or deployment. We propose an interaction-efficient active learning framework in which a dueling bandit algorithm iteratively selects pairs of LLMs, collects user feedback about their responses, and updates its belief about the user's latent preferences. We introduce a novel belief-aware upper confidence bound strategy that balances exploration of the model pool with exploitation of inferred preferences, enabling efficient alignment between user needs and LLM capabilities under user-specified cost and time budgets. Through diverse experiments on LLMs and human studies, we experimentally verify that our model can efficiently match well-aligned LLMs to users at a lower cost.
Son Nguyen, Xinyuan Liu, Ransalu Senanayake
School of Computing and Augmented Intelligence, Arizona State University, Tempe, United States of America
Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on the models producing measurably different conversational behaviors when given the same input. Prior offline studies recommend drawing one model per family for behavioral diversity, because LLMs prefer outputs from their own family when rating one another in isolation. Whether the same family label predicts behavior in interactive multi-LLM systems, the setting that real deployed systems use, has not been tested. We study this with a 940,000-chain 11-checkpoint corpus and a 1.6M-chain same-base Llama factorial. On our validated headline metric, hedging, a reasoning-distilled Llama checkpoint shifts by 18% depending on which same-base partner it replies to, more than any cross-family hedging gap in the controlled subset. Qwen, closed-API, and runtime checks suggest the pattern is not isolated, while repair and challenge analyses remain exploratory because their surface-cue detectors are weaker. Overall, the results identify post-training recipe as a first-class axis for multi-LLM panel composition and show that model family alone is an incomplete proxy for conversational diversity.
Luyang Zhang, Jialu Wang, Fei Xue +1
Carnegie Mellon University · University of California, Santa Cruz · Independent Researcher