Decoupled Multi-Agent Orchestration
Organizations: National University of Singapore
Abstract
Learned orchestration can automatically construct effective language-model multi-agent systems, but existing approaches couple planning to fixed worker pools and train decomposition and collaboration from the same terminal outcome, limiting transfer and obscuring credit assignment. We introduce DeOrch, which separates worker-agnostic planning from concrete worker selection. Its two-stage planner first decomposes the task without worker information, then chooses collaboration operations using compact, worker-identity-free matchability feedback from the pool, enabling conditional credit assignment to decomposition and collaboration decisions. A lightweight matcher estimates worker suitability from behavior on a fixed probe set and adapts online with a contextual bandit, allowing new workers to be incorporated without retraining the planner or matcher. Across diverse in- and out-of-distribution tasks, DeOrch outperforms prior automatic MAS orchestration methods with fewer worker calls than competing learned orchestrators, remains effective when transferred to an entirely unseen worker pool without retraining, and shows consistent gains from both components.
Figures & tables
| In-Distribution | Out-of-Distribution | Overall | ||||||||
| Method | MuSiQue | TP | Omni | LCB | BBEH | NP | PB-M | GAIA | Avg. | Calls |
| Best Single Agent | 72.8 | 13.6 | 42.5 | 69.9 | 38.4 | 51.1 | 19.4 | 6.1 | 39.2 | 1.00 |
| Independent | 72.0 | 9.2 | 47.4 | 76.4 | 45.2 | 63.1 | 25.9 | 28.5 | 46.0 | 7.00 |
| Frontier-Planner | 59.1 | 3.5 | 38.9 | 71.0 | 37.4 | 45.5 | 9.1 | 28.5 | 36.6 | 4.67 |
| Untrained Planner | 60.2 | 3.3 | 38.1 | 70.4 | 31.8 | 50.4 | 25.2 | 26.7 | 38.3 | 4.92 |
| Conductor | 70.3 | 9.9 | 41.2 | 75.0 | 37.1 | 52.9 | 38.0 | 27.9 | 44.0 | 6.16 |
| In-Distribution | Out-of-Distribution | Overall | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | MuSiQue | TP | Omni | LCB | BBEH | NP | PB-M | GAIA | Avg. |
| Best Single Agent | 69.0 | 6.2 | 45.9 | 75.9 | 47.7 | 40.3 | 29.0 | 12.1 | 40.8 |
| Independent | 72.4 | 5.0 | 50.6 | 81.2 | 48.4 | 60.4 | 41.6 | 20.0 | 47.5 |
| Frontier-Planner | 61.4 | 3.1 | 44.4 | 76.3 | 46.8 | 52.2 | 18.1 | 21.8 | 40.5 |
| Conductor | 63.7 | 4.6 | 45.5 | 75.5 | 46.2 | 48.8 | 23.7 | 18.8 | 40.9 |
| DeOrch | 75.5 | 11.3 | 52.3 | 83.0 | 49.7 | 62.3 | 44.3 | 23.0 | 50.2 |
| Variant | ID Avg. | OOD Avg. | Avg. | Calls |
|---|---|---|---|---|
| Joint Planning | 43.2 | 50.2 | 48.5 | 6.29 |
| Two-Stage | 43.8 | 50.9 | 49.2 | 7.33 |
| + Matchability | 44.3 | 51.4 | 49.6 | 7.52 |
| + Conditional Credit | 47.7 | 53.4 | 51.9 | 7.78 |
| Decomposition + Single | 46.1 | 51.8 | 50.4 | 1.68 |
| DeOrch | 47.4 | 52.9 | 51.5 | 5.35 |
| Method | ID Avg. | LOSO-OOD Avg. | Calls |
|---|---|---|---|
| Random | 59.44 | 69.73 | 1 |
| Best Single Worker | 68.33 | 75.26 | 1 |
| Majority vote | 67.92 | 74.59 | 6 |
| Probe Matcher w/o Density | 69.70 | 75.36 | 1 |
| Probe Matcher | 70.24 | 75.45 | 1 |
| Oracle Single | 82.68 | 87.17 | 1 |
| Adaptation Examples | 0 | 3K | 6K | 9K | 12K | 15K |
|---|---|---|---|---|---|---|
| Overall Avg. | 51.3 | 51.7 | 52.1 | 52.1 | 52.6 | 52.9 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Usage | Dataset | Primary task type | # Examples | Sampling |
|---|---|---|---|---|
| Probe set | APPS | Code generation | 4,962 | Full |
| BFCL | Tool / function use | 3,151 | Full | |
| Big-Math | Mathematical reasoning | 6,000 | Sampled | |
| FRAMES | Knowledge-intensive reasoning | 824 | Full | |
| MBPP+ | Code generation | 378 | Full | |
| SuperGPQA | General knowledge / reasoning | 26,529 | Full |
| Benchmark | Total | Adaptation | Held-out |
|---|---|---|---|
| MuSiQue | 2,000 | 1,750 | 250 |
| TravelPlanner | 1,000 | 800 | 200 |
| Omni-MATH-2 | 4,428 | 4,128 | 300 |
| LiveCodeBench v6 | 611 | 411 | 200 |
| BBEH | 4,520 | 4,220 | 300 |
| NaturalPlan | 3,600 | 3,300 | 300 |
| Selection Rule | ID Avg. | OOD Avg. | Overall Avg. |
|---|---|---|---|
| Top- Suitability | 45.9 | 52.3 | 50.7 |
| Complementarity-Aware selection | 47.4 | 52.9 | 51.5 |
| Method | Avg. | Calls | Tokens / Query |
|---|---|---|---|
| Best Single Agent | 39.2 | 1.00 | 1859 |
| Best Single + Refinement | 42.4 | 4.00 | 14106 |
| Independent | 46.0 | 7.00 | 25859 |
| Frontier-Planner | 36.6 | 4.67 | 11479 |
| Untrained Planner | 38.3 | 4.92 | 12867 |
| Conductor | 44.0 | 6.16 | 22458 |
| Variant | ID Avg. | OOD Avg. | Avg. | Calls |
|---|---|---|---|---|
| Counterfactual Credit | 43.1 | 51.2 | 49.2 | 5.17 |
| Two-Stage Conditional Credit ( DeOrch ) | 47.4 | 52.9 | 51.5 | 5.35 |