Learned orchestration can automatically construct effective language-model multi-agent systems, but existing approaches couple planning to fixed worker pools and train decomposition and collaboration from the same terminal outcome, limiting transfer and obscuring credit assignment. We introduce DeOrch, which separates worker-agnostic planning from concrete worker selection. Its two-stage planner first decomposes the task without worker information, then chooses collaboration operations using compact, worker-identity-free matchability feedback from the pool, enabling conditional credit assignment to decomposition and collaboration decisions. A lightweight matcher estimates worker suitability from behavior on a fixed probe set and adapts online with a contextual bandit, allowing new workers to be incorporated without retraining the planner or matcher. Across diverse in- and out-of-distribution tasks, DeOrch outperforms prior automatic MAS orchestration methods with fewer worker calls than competing learned orchestrators, remains effective when transferred to an entirely unseen worker pool without retraining, and shows consistent gains from both components.
Figures & tables
Figure 1: Overview of DeOrch .
In-Distribution
Out-of-Distribution
Overall
Method
MuSiQue
TP
Omni
LCB
BBEH
NP
PB-M
GAIA
Avg.
Calls ↓
Best Single Agent
72.8
13.6
42.5
69.9
38.4
51.1
19.4
6.1
39.2
1.00
Independent
72.0
9.2
47.4
76.4
45.2
63.1
25.9
28.5
46.0
7.00
Frontier-Planner
59.1
3.5
38.9
71.0
37.4
45.5
9.1
28.5
36.6
4.67
Untrained Planner
60.2
3.3
38.1
70.4
31.8
50.4
25.2
26.7
38.3
4.92
Conductor
70.3
9.9
41.2
75.0
37.1
52.9
38.0
27.9
44.0
6.16
Table 1: Main results across in-distribution and out-of-distribution benchmarks.
In-Distribution
Out-of-Distribution
Overall
Method
MuSiQue
TP
Omni
LCB
BBEH
NP
PB-M
GAIA
Avg.
Best Single Agent
69.0
6.2
45.9
75.9
47.7
40.3
29.0
12.1
40.8
Independent
72.4
5.0
50.6
81.2
48.4
60.4
41.6
20.0
47.5
Frontier-Planner
61.4
3.1
44.4
76.3
46.8
52.2
18.1
21.8
40.5
Conductor
63.7
4.6
45.5
75.5
46.2
48.8
23.7
18.8
40.9
DeOrch
75.5
11.3
52.3
83.0
49.7
62.3
44.3
23.0
50.2
Table 2: Performance under an unseen worker pool.
Variant
ID Avg.
OOD Avg.
Avg.
Calls ↓
Joint Planning
43.2
50.2
48.5
6.29
Two-Stage
43.8
50.9
49.2
7.33
+ Matchability
44.3
51.4
49.6
7.52
+ Conditional Credit
47.7
53.4
51.9
7.78
Decomposition + Single
46.1
51.8
50.4
1.68
DeOrch
47.4
52.9
51.5
5.35
Table 3: Ablation of the planner.
Method
ID Avg. ↑
LOSO-OOD Avg. ↑
Calls ↓
Random
59.44
69.73
1
Best Single Worker
68.33
75.26
1
Majority vote
67.92
74.59
6
Probe Matcher w/o Density
69.70
75.36
1
Probe Matcher
70.24
75.45
1
Oracle Single
82.68
87.17
1
Table 4: Offline evaluation of worker matching.
Adaptation Examples
0
3K
6K
9K
12K
15K
Overall Avg.
51.3
51.7
52.1
52.1
52.6
52.9
Table 5: Online adaptation with increasing deployment feedback.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Usage
Dataset
Primary task type
# Examples
Sampling
Probe set
APPS
Code generation
4,962
Full
BFCL
Tool / function use
3,151
Full
Big-Math
Mathematical reasoning
6,000
Sampled
FRAMES
Knowledge-intensive reasoning
824
Full
MBPP+
Code generation
378
Full
SuperGPQA
General knowledge / reasoning
26,529
Full
Appendix
Table 6: Data sources used for probe construction, planner post-training, and evaluation.
Benchmark
Total
Adaptation
Held-out
MuSiQue
2,000
1,750
250
TravelPlanner
1,000
800
200
Omni-MATH-2
4,428
4,128
300
LiveCodeBench v6
611
411
200
BBEH
4,520
4,220
300
NaturalPlan
3,600
3,300
300
Appendix
Table 7: Dataset split for the online adaptation experiment. The held-out examples are used only for evaluation and never for posterior updates.
Selection Rule
ID Avg. ↑
OOD Avg. ↑
Overall Avg. ↑
Top- n Suitability
45.9
52.3
50.7
Complementarity-Aware selection
47.4
52.9
51.5
Appendix
Table 8: Inference-time ablation of multi-worker selection.
Method
Avg. ↑
Calls ↓
Tokens / Query ↓
Best Single Agent
39.2
1.00
1859
Best Single + Refinement
42.4
4.00
14106
Independent
46.0
7.00
25859
Frontier-Planner
36.6
4.67
11479
Untrained Planner
38.3
4.92
12867
Conductor
44.0
6.16
22458
Appendix
Table 9: Inference cost across the eight evaluation benchmarks.
Variant
ID Avg.
OOD Avg.
Avg.
Calls ↓
Counterfactual Credit
43.1
51.2
49.2
5.17
Two-Stage Conditional Credit ( DeOrch )
47.4
52.9
51.5
5.35
Appendix
Table 10: Comparison of credit-assignment formulations.
Large language model (LLM) multi-agent systems typically rely on rigid orchestration, committing either to flat per-query routing or to hand-engineered task decomposition, so decomposition depth, worker choice, and inference budget are not jointly optimized under one objective. We introduce Uno-Orchestra, a unified orchestration policy that selectively decomposes a task and dispatches each subtask to an admissible (model, primitive) pair, with both decisions learned together from curated RL trajectories grounded in real worker interactions. Against 22 baselines on a 13-benchmark suite spanning math, code, knowledge, long-context, and agentic tool-use, Uno-Orchestra reaches 77.0% macro pass@1, roughly 16% above the strongest workflow baseline, at roughly an order of magnitude lower per-query cost, advancing the accuracy-efficiency frontier of selective delegation.
Multi-Agent Systems (MAS) built on Large Language Models (LLMs) require effective orchestration to coordinate specialized agents, yet training such orchestrators is hindered by limited supervision and high computational cost. We propose Orchestration Reward Modeling (OrchRM), a self-supervised framework for evaluating orchestration quality without human annotations. OrchRM leverages intermediate artifacts from multi-agent executions to construct win-lose pairs for Bradley-Terry reward model training. Unlike existing MAS test-time scaling and orchestrator training frameworks that rely on costly sub-agent rollouts, OrchRM operates directly at the orchestration level, enabling efficient and high-performing reward-guided orchestrator training and MAS test-time scaling. OrchRM improves training efficiency by up to 10x in token usage while improving MAS test-time scaling performance by up to 8% in accuracy. These gains consistently transfer across multiple domains, including mathematical reasoning, web-based question answering, and multi-hop reasoning, demonstrating orchestration-level reward modeling as a scalable direction for robust multi-agent orchestration. Code will be available at https://github.com/Wang-ML-Lab/OrchRM.
King Yeung Tsang, Zihao Zhao, Vishal Venkataramani +5
Multi-agent systems (MAS) demonstrate clear advantages in tackling complex problems by coordinating diverse agents and external tools. However, most existing orchestration methods rely on static workflows or serial agent scheduling, and are further constrained by heterogeneous interface protocols between tools and agents. This leads to high system complexity and poor extensibility. To mitigate these issues, we propose Agent-as-Tool, a unified parallel orchestration paradigm that abstracts both agents and tools into a standardized, learnable action space with protocol normalization and explicit state feedback. Building on this paradigm, we train a lightweight orchestrator, ParaManager, which decouples planning decisions from subtask solving, enabling state-aware parallel subtask decomposition, delegation, and asynchronous execution. For training, we adopt a two-stage ParaManager training pipeline. It improves robustness by incorporating supervised fine-tuning (SFT) trajectories equipped with recovery mechanisms, and further applies reinforcement learning (RL) to achieve an optimal balance among task success, protocol compliance, diversity, and reasoning efficiency. Experiments show that ParaManager achieves strong performance across multiple benchmarks and exhibits robust generalization under unseen model pools.
Wenzhen Yuan, Wutao Xiong, Fanchen Yu +7
Shanghai Jiao Tong University · Sichuan University · Shanghai AI Lab +2