LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pursues but does not reveal. Inferring such hidden preferences from behavior has been a subject of long-standing research in game theory and multi-agent systems. The core challenge lies in maintaining a belief over every agent's preference and updating it from the agents' observed actions. Existing LLM orchestrators carry that belief as prompt text with no explicit update rule. This lets early errors persist and propagate rather than be corrected. We therefore propose \textbf{HARP} (Heterogeneous-preference Agent oRchestration via Preference inference), a novel framework that moves the belief out of the prompt. Specifically, HARP maintains one numeric posterior per agent over a finite set of candidate preferences and updates it in closed form by Bayes' rule. The language model supplies only actions and per-candidate likelihoods, so estimation is decoupled from its reasoning. We prove that HARP attains the same O~(K) Bayesian regret as explicit joint inference when the factorization is exact. Furthermore, HARP\textsuperscript{+} augments planning with a bonus for actions that distinguish the candidates, so inference continues even when the optimal action is uninformative. Empirical results on three substrates, ranging from payoffs the preferences fully determine, through payoffs that depend on more than them, to scales where explicit joint inference is infeasible, demonstrate that HARP\textsuperscript{+} is the strongest non-oracle method across the class our theory identifies.
Figures & tables
Figure 1: Overview of HARP. (a) An orchestrator dispatches a joint assignment u=(u1,…,un) to LLM agents and sees states, actions and rewards, never private personas θi⋆ ; prompt-text beliefs let errors persist. (b) HARP keeps per-agent numeric posteriors outside the prompt: (1) prior; (2) sample θ^i∼μki ( ▼ ); (3) plan with P (HARP + adds a persona-discrimination bonus) and dispatch u ; (4) the offline-calibrated scorer q evaluates a under every candidate persona; (5) Bayes-update each posterior without LLM calls; repeat. (c) Under the structural conditions (TI), (RL), (PF), per-agent posteriors reconstruct the joint posterior exactly, storing n∣Θi∣ , not ∣Θi∣n , entries.
Figure 2: Cumulative Bayesian regret on HP-SPGG at K=20 for four LLM backbones, 10 common seeds. Error bars are the standard error of the mean.
Figure 3: HP-SPGG scaling and component ablation (Llama-4-Maverick, ∣Θi∣=4 , K=20 , 10 common seeds). Top : cumulative regret against agent count, where HARP and Joint-PSRL sit at the same regret level at every n . Middle : persona storage against agent count, annotated with the joint-to-factored ratio. Bottom : cumulative regret for HARP + with one belief component removed, labelled ( − bonus, − update, − identity) and by decentralized execution ( − dispatch), on a log scale with the oracle at zero.
Figure 4: Focal-agent payoff on the Concordia substrates. (a) Pub Coordination against oracle_joint . (b) Haggling against oracle_focal . The full set of configurations, with HARP + ’s margin, is in Figure 25 .
Figure 5: MaaSSim dispatch over 10 common environment seeds. (a) Realized utility across conflict strength λ . (b) Oracle regret in reject-penalty units. (c) Realized utility for each source of the belief the dispatcher acts on. Error bars are seed-level standard errors in (a) and (c), and oracle and policy standard errors propagated in quadrature in (b).
Figure 6: Iterated Concordia, 5 common seeds. (a) Paired HARP minus Joint-PSRL regret for each geometry and pooled, where negative values favor HARP. Whiskers are t -based 95% intervals with 4 degrees of freedom. (b) Update value, the paired excess of the no-update variant over HARP + , against persona decision value. Bars are the standard error of the mean.
Figure 7: MaaSSim parity and mechanism. (a) Relative reduction in oracle regret from the joint persona posterior, where the top row pools six independent-prior cells. One marker per row, since HARP + reproduces Joint-PSRL on every round. (b) Per-event update time, with the joint table not run beyond n=4 . (c) Realized utility against the accuracy of the belief the dispatcher acts on, oracle dashed in red. (d) Per-event posterior gain by event type. Error bars are the standard error of the mean.
Table 1: Research-question reading guide. The appendix mirrors the main-text organization by research question.
Backbone
HARP +
HARP
Joint-PSRL
LLM-PSRL
Best coordination
Coord./best HARP
DeepSeek-V3.2
0.360±0.142
0.985±0.298
0.955±0.203
18.270±5.141
ECON-BNE 2.550±1.049
7.08×
GPT-5.4-nano
0.159±0.038
0.544±0.150
0.391±0.100
2.193±0.714
ECON-BNE 6.960±1.376
43.77×
Kimi-K2.6
0.325±0.088
0.547±0.116
0.642±0.159
6.701±1.325
ECON-BNE 2.957±1.366
9.10×
Llama-4-Maverick
0.430±0.217
1.202±0.212
1.276±0.228
13.510±3.431
A-ToM-0 0.700±0.396
1.63×
Appendix
Table 2: Environment-matched HP-SPGG control ( K=20 , 10 common seeds). Values are cumulative regret, mean ± SEM. Ratio is the strongest fixed LLM-coordination baseline divided by the lower-regret HARP-family member (HARP + on all four backbones).
Figure 8: SOTOPIA-Hard descriptive results for three private-goal families: craigslist_bargains ( n=80 ), revenge_plot ( n=20 ), and donate_funds ( n=20 ). (a) End-of-dialogue focal score across four backbones (SEM); Δ compares HARP + with the next-best LLM-coordination baseline. (b) Prefix-only per-turn trajectory (descriptive; the historical posterior did not update, so this panel is not concentration evidence); red dashed is oracle_policy (Appendix A.4 ).
Family
n
HARP +
Best
Δ
craigslist_bargains
80
3.04
2.98
+0.06
revenge_plot
20
2.93
2.83
+0.10
donate_funds
20
3.26
3.24
+0.02
Appendix
Table 3: SOTOPIA-Hard private-goal families: HARP + vs. strongest LLM-coordination baseline per family, averaged across four backbones. Best baseline per family: A-ToM-1 for craigslist_bargains and donate_funds ; ECON-BNE for revenge_plot .
Backbone
HARP +
A-ToM-1
ECON-BNE
llm_belief
llm_greedy
Rank
Δ
DeepSeek-V3.2
3.248
3.261
3.313
3.195
3.193
3
−0.065
GPT-5.4-nano
2.981
2.918
2.843
3.000
2.842
2
−0.019
Kimi-K2.6
2.807
2.733
2.876
2.906
2.761
3
−0.099
Llama-Maverick
3.308
3.359
3.319
3.340
3.165
4
−0.051
Average
3.086
3.068
3.088
3.110
2.990
—
−0.024
Appendix
Table 4: SOTOPIA-Hard aggregate mean focal score per (backbone × baseline), 70 episodes per cell. Bold marks the highest non-oracle score per backbone; HARP + rank and gap to best alternative are in the last two columns. The two oracle rows (Appendix A.4 ) receive the privileged one-hot opponent persona; oracle_policy additionally selects the best of K=5 independent episodes by the same seven-dimensional judge.
Figure 9: Corrected SOTOPIA posterior diagnostic on GPT-5.4-nano. (a) Recurrent mass on the profile-derived proxy persona (not a native SOTOPIA truth label); error bars are episode-row SEM. (b) Focal score under deterministic intent-menu corruption; provider and judge generations are not pathwise seed-controlled, so this is a sensitivity bound rather than a paired causal dose response.
Family
Corrected score
Δ vs best
Craigslist
2.665±0.074
−0.095
Donate
3.229±0.121
−0.157
Revenge
2.743±0.102
−0.457
Appendix
Table 5: Corrected p=0 SOTOPIA branch decision. Historical best is the retained GPT-nano comparator, not a contemporaneous paired rerun.
Figure 10: Corrected SOTOPIA component variants on GPT-5.4-nano (same 30 cases, four replicates; episode-row SEM). HARP + does not significantly exceed either corrected control on any family. Provider generations are not pathwise seed-controlled.
Figure 11: C1 analytic-tier nine-baseline n -scaling on HP-SPGG ( n∈{3,4,5,6} , K=20 , 5 seeds, β=0.25 ). The HARP family stays under 1.0 cumulative regret across the sweep while type-agnostic IQL and Random grow with n . Persona-storage savings are 5.3× , 16× , 51× , and 171× at n=3,4,5,6 .
Figure 12: C1 live-LLM nine-baseline n -scaling on HP-SPGG c19 ( K=20 , 5 common environment seeds, β=0.25 ). Left: DeepSeek-V3.2. Right: Llama-4-Maverick. The HARP family remains separated from type-agnostic baselines on both backbones; external A-ToM-0/1/2 and ECON-BNE references are shown at n=3 .
Storage
analytic kernel
DeepSeek-V3.2
n
savings
HARP
J-PSRL-U
HARP
J-PSRL-U
3
5.3 ×
0.12
0.17
0.57
0.38
4
16 ×
0.05
0.11
0.80
0.51
5
51 ×
0.60
0.63
1.02
0.85
Appendix
Table 6: Factored-vs-joint storage comparison on HP-SPGG ( m=∣Θi∣=4 , K=100 , 10 common environment seeds, β=0.25 ). HARP and Joint-PSRL-Uniform share likelihood, exhaustive planner, and seeds; only the posterior representation differs. Values are cumulative regret.
n
λ
Gap
95% CI
Max TV
Entries
2
0
+1.85
[+0.19,+3.51]
7.8×10−16
256 vs 32
2
1
+0.11
[−2.87,+3.10]
7.8×10−16
256 vs 32
3
0
+1.33
[−0.71,+3.37]
2.5×10−15
4096 vs 48
3
1
−1.59
[−5.13,+1.95]
2.5×10−15
4096 vs 48
4
0
−1.38
[−3.90,+1.15]
2.5×10−14
65536 vs 64
4
1
−0.37
[−4.08,+3.34]
2.5×10−14
65536 vs 64
Appendix
Table 7: E-E paired utility gap (joint − factored; 10 common environment seeds, mean and 95% CI) and posterior identity.
n
storage ratio
joint update
HARP +
HARP = Joint
PSRL-NoType
2
8×
7.5μ s
0.009±0.009
0.090±0.036
2.55±0.50
3
85×
12.1μ s
0.032±0.022
0.132±0.037
3.00±0.42
4
1,024×
72.9μ s
0.036±0.024
0.148±0.034
3.75±0.42
5
13,107×
3.2 ms
0.037±0.022
0.243±0.052
4.15±0.55
6
174,763×
68.8 ms
0.045±0.023
0.318±0.062
5.41±0.72
7
2.4×106
infeasible
0.036±0.024
0.380±0.063 (HARP)
6.01±0.52
Appendix
Table 8: Frontier sweep at ∣Θi∣=16 (mean ± SEM over ten seeds, K=50 ). HARP and Joint-PSRL are pathwise identical wherever both run. Joint-PSRL leaves the sweep at n=7 by the 1 s update rule and at n=8 by the 4 GB rule.
Figure 13: C4 response-locality violation ( K=100 , 10 common environment seeds). Left: HARP and correctly specified Joint-PSRL regret trajectories at α=4 . Right: marginal TV between HARP’s factored posterior and a shadow exact-joint posterior on the same trajectory. Top is analytic; bottom uses all 324 live DeepSeek persona/player/action cells. Posterior coupling increases, but no regret disadvantage is detectable on this geometry.
Tier
α
HARP
Joint
Gap
Analytic
0
.001±.001
.001±.001
.000±.000
Analytic
1
.155±.065
.206±.087
−.051±.051
Analytic
4
.199±.057
.213±.057
−.013±.013
DeepSeek live
0
1.130±.154
1.130±.154
.000±.000
DeepSeek live
1
.583±.115
.575±.116
.008±.008
DeepSeek live
4
.517±.131
.466±.108
.051±.107
Appendix
Table 9: C4 RL-violation cumulative regret at K=100 ; Gap is the common-environment-seed HARP-minus-Joint difference.
Figure 14: C4 prior swap on HP-SPGG ( n=3 , ∣Θ∣=4 , ∣Θ∣n=64 joint profiles, K=20 , 10 common environment seeds, β=0.25 , analytic calibration). Left: symmetric Dirichlet sweep on the 64 -simplex; each seed draws p∼Dir(α164) and HARP inherits the induced marginals while the explicit joint baseline uses p . Right: a shared-type prior, uniform over the four same-type profiles, leaves marginals uniform while maximally coupling them. The HARP family stays within 0.6 cumulative regret on every retained cell.
Substrate
n
HARP
J-PSRL-A
HARP / Aware
analytic kernel
3
0.28
0.22
1.29 ×
4
0.00
0.32
0.00 ×
5
0.31
0.39
0.79 ×
DeepSeek-V3.2
3
1.79
1.23
1.46 ×
4
1.31
1.84
0.71 ×
5
2.11
2.61
0.81 ×
Appendix
Table 10: True shared-type DGP robustness on HP-SPGG across one analytic kernel and four live LLM backbones ( m=∣Θi∣=4 , K=100 , 10 common environment seeds, β=0.25 ). True types are drawn from the m same-type joint profiles, so Joint-PSRL-Aware’s strict prior is well-specified and HARP’s factored prior is mis-specified relative to the coupled truth. HARP remains within 1.5× of Aware on 14 of 15 cells (Kimi-K2.6 n=4 at 1.52× ); the GPT-5.4-nano n=4 cell where Aware reaches 0 is marked with a dash. HARP is lower-regret than Aware on 9 of 15 cells.
Figure 15: Rate and burn-in evidence. (a) HARP and HARP + plateau at <0.05 regret per episode within K≈5 on four HP-SPGG backbones, while PSRL-NoType grows linearly and reaches an 11 – 30× cumulative gap at K=20 . (b) Vanilla-HARP concentration time Kconc0.9 predicts the benefit of the β=0.25 discrimination bonus across the four backbones.
Figure 16: C2 analytic-kernel cross-check: HARP + on HP-SPGG c19 remains within 0.05 regret of the privileged analytic-kernel reference across all four backbones.
Method
Cumulative regret
Late regret
HARP
0.312±0.026
0.0041±0.0012
HARP +
0.352±0.035
0.0020±0.0017
Joint-PSRL
0.329±0.030
0.0049±0.0012
PSRL-NoType
1.361±0.172
0.0764±0.0139
Appendix
Table 11: Iterated Concordia-derived aggregate at K=20 . Late is mean instant regret over episodes 11 – 20 .
Figure 17: C3 scaling of HARP and HARP + on HP-SPGG. Left varies ∣Θi∣ at n=3 ; right varies n at ∣Θi∣=4 . Curves report K=20 cumulative regret over 10 seeds on Llama-4-Maverick with β=0.25 .
Figure 18: C3 HP-SPGG β -sweep at K=20 . The frozen value β=0.25 is the smallest point in the no-degradation interval [0.25,0.75] ; it leaves three saturated backbones unchanged and reduces Llama-Maverick regret from 0.97 to 0.31 .
HARP
HARP +
Δ
Concordia (single-shot)
Pub Coordination
1.264
1.264
+0.000
Haggling (single-item)
3.661
3.661
+0.000
Haggling (multi-item)
5.429
5.429
+0.000
MaaSSim replay (E-F, 10 seeds, utility)
Dispatch utility
27.61
27.36
−0.25
Appendix
Table 12: C3 direct type-discrimination bonus ablation. HARP and HARP + share likelihood, posterior, sampling, and solver; only β changes. Concordia is single-shot, so the multi-step concentration mechanism is inactive. MaaSSim rows are E-F fixed-state replay with frozen β=0.25 ; the bonus changes 4 of 406 assignments and the paired gap covers zero. Historical SOTOPIA rows are descriptive because the posterior in those runs remained at its prior.
Figure 19: Price of decentralization on Concordia london_mini Pub Coordination ( 5 players, 2 focal, 5 scenes per cell). Centralized (planner LLM, per-agent type-guess) vs. decentralized (ToM-prompted, greedy), same LLM. (a) Focal payoff. (b) Coordination. (c) Social welfare. (d) Cumulative regret over K=5 scenes (bands: SE over backbones). Centralized modes cluster near oracle_joint ; decentralized modes: −0.2 – 0.3 payoff, −0.2 – 0.4 coordination, −1 – 2 welfare units, 3 – 4× cumulative regret.
Figure 20: HARP driver-belief concentration on MaaSSim ( 10 seeds, thin lines are per-seed means, bands are SEM over (seed, driver) pairs). From accept/reject observations alone, marginal rule accuracy and exact-type posterior mass rise well above their chance references within a handful of observations of each driver.
Variant
Belief source
Utility
Rejects
Accept
Rule acc.
Nearest
none
11.03
26.0
0.634
—
Random
none
−33.94
30.6
0.563
—
HARP-prior
uniform
11.11
25.1
0.635
0.500
HARP-shuffled
wrong driver
6.87
26.5
0.609
0.521
HARP
learned
27.61
18.5
0.740
0.720
Oracle
true persona
38.92
13.2
0.822
1.000
Appendix
Table 13: MaaSSim persona-mechanism replay ( 10 seeds, fixed queue snapshots and persona maps, shared assignment objective). Utility is the realized dispatch utility defined above (seed-level SEM 9.6 – 11.7 ); Rejects is mean driver rejects; Accept is the driver acceptance rate; Rule acc. is the marginal accuracy of the belief on the true driver rule.
Figure 21: Operating points of the mechanism-replay variants ( 10 seeds, SEM bars on both axes). HARP trades pickup wait for fewer driver rejects, moving from the uniform-prior operating point toward the oracle (dashed arrow), while the shuffled control stays at the heuristic wait level with more rejects. This is why Table 13 supports a persona-mechanism claim rather than a wait-minimization claim.
Figure 22: MaaSSim persona-mechanism replay. All variants share the same fixed queue snapshots, persona maps, and assignment objective, and only the driver-belief source changes. The learned posterior (HARP) closes 59.3% of the prior-to-oracle utility gap, while the shuffled-posterior control falls below the uniform prior, showing that correct driver-identity attachment is necessary for the gain.
Scenario
HARP
Best prompt
Gap
Oracle
Reject ( λ=0 )
18.37
PSRL: 13.77
+4.60
36.44
Mid ( λ=.5 )
8.96
belief: −22.80
+31.75
30.70
Full ( λ=1 )
8.79
belief: −30.90
+39.69
22.44
Appendix
Table 14: MaaSSim LLM dispatch conflict-strength sweep ( gpt-5.4-mini , 10 seeds ×20 snapshots, driver-reject penalty 5.0 , parse rate 1.000 ). Best prompt is the strongest pure prompt baseline per level; Gap is HARP minus best prompt. Outcomes are measured reruns at λ∈{0,0.5,1} , not interpolated results.
Figure 23: Single-episode conflict-offer replay dynamics ( λ=1 ) for HARP, LLM-PSRL, A-ToM-1, and ECON-BNE on a common seed: vehicle and request traces with cumulative utility, served requests, and driver rejects per decision step.
Figure 24: Unit validation for the main-text regret panel. Oracle regret in reject-penalty units against excess driver rejects relative to the oracle, for six methods across three measured replay scenarios ( 10 seeds).
Figure 25: Full Concordia comparison across all 18 configurations. Pub Coordination uses oracle_joint ; Haggling uses the exact full-information focal ceiling oracle_focal . Rows are sorted by HARP + ’s margin over the strongest baseline.
Figure 26: Focal-vs-joint Haggling frontier. βobj=0 is oracle_joint (J); βobj=1 is oracle_focal (F). Increasing focal payoff can reduce minimum participant surplus, explaining why a focal policy may exceed the joint oracle’s focal score without violating an upper bound.
Benchmark
Payoff form
Action space
Oracle type
HP-SPGG
analytic/pinned tensor
finite, 125 joint profiles
strict exhaustive argmax
Concordia
analytic, closed-form
finite menu
strict joint/focal argmax
SOTOPIA-Hard
7-dimensional LLM judge
open-ended utterance
belief-only + best-of- K critic
Appendix
Table 15: Oracle construction across the three LLM benchmarks.
Main-text claim
Formal statement
Proof location
Exact factorization and persona-belief complexity
Thm. B.11 , Cor. B.12
§ B.3
Neither TI nor RL can be dropped
Prop. B.14 + controlled diagnostic
§ B.4 , § A.2
O~(K) regret bound for HARP, for the team and per agent
Thm. B.15 , Cor. B.16
§ B.5
Rate separation vs PSRL-NoType
Thm. B.35
§ B.9
HARP + burn-in improvement
Thm. B.37
§ B.10
Robustness to scorer error
Prop. B.26
§ B.6
Appendix
Table 16: Claim-to-theorem map. Each main-text statement is keyed to the formal theorem and the appendix subsection containing the proof.
n
2
4
8
16
32
64
burn-in
95
113
143
175
207
236
Appendix
Table 17: Population phase of the preregistered contraction study ( m=8 , H=4 , 500 seeds per cell, zero censoring). All-agent burn-in grows with logn , not linearly in n .
Role
HP-SPGG
Concordia
MaaSSim
Public state
Round index and the previous contribution profile
None in the native format, where each scene is one decision
Queue snapshot
Assignment u
One contribution level per player, 125 joint profiles
A joint choice from the finite scene menu
A one-to-one assignment of offers to drivers
Emitted action ai
Satisfaction score of player i with the assigned profile, also its reward
The players’ actions in the native one-shot format, the normalized payoff in the iterated variant
Accept or reject
Agents
The backbone under a persona prompt. Its scores are collected offline in a tensor with one entry per persona, player, and profile. During the loop a response is the tensor entry plus Gaussian noise of scale 0.08 , clipped to [0,1] . The analytic tier replaces the backbone’s scores with a closed-form payoff
Scene players. Payoffs come from the analytic table payoff_for_case , and no scripted non-player policy is executed
Simulator drivers with rule-based acceptance, 16 rule combinations
Scorer q
Gaussian around the same tensor, 1,500 judge cells per backbone, all offline
Computed from the analytic payoff, Gaussian after normalization in the iterated variant
Per-driver accept or reject likelihood of each rule combination
Planner P
Backward induction with an exhaustive argmax over the 125 profiles
Enumeration of the finite one-step menu
Exhaustive assignment objective in the parity and mechanism replays. In the sweep of Figure 5 (a,b), a gpt-5.4-mini dispatcher selects from the legal menu with HARP’s assignment scores in its context. That dispatcher lies outside (DP)
Appendix
Table 18: Roles in each substrate. The last row counts the language-model calls that HARP issues while the loop runs.
Large language models (LLMs) provide a flexible foundation for multi-agent systems, but their effectiveness and computational cost depend critically on orchestration design. Across different tasks, role design, capacity assignment, and dependency construction jointly affect both solution quality and execution efficiency. Existing approaches automate parts of this design process, yet they often optimize these decisions partially or sequentially, and rely on execution-level feedback that provides limited credit assignment for local orchestration decisions. We propose LEMON (Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning), an LLM-based orchestrator that learns to design efficient multi-agent orchestration. Given a task, LEMON designs a unified orchestration specification that composes customized agent duties, capacity levels, and dependency relations. To train the orchestrator, we augment the orchestration-level Group Relative Policy Optimization (GRPO) objective with a localized counterfactual credit signal that edits role, capacity, or dependency fields and applies the resulting reward contrast only to the edited spans. Experiments on six reasoning and coding benchmarks, including MMLU, GSM8K, AQuA, MultiArith, SVAMP, and HumanEval, show that LEMON achieves the best average performance among evaluated multi-agent orchestration methods while improving the accuracy-token trade-off.
LLMs excel at predictive tasks and complex reasoning tasks, but many high-value deployments rely on decisions under uncertainty, for example, which tool to call, which expert to consult, or how many resources to invest. While the usefulness and feasibility of Bayesian approaches remain unclear for LLM inference, this position paper argues that the control layer of an agentic AI system (that orchestrates LLMs and tools) is a clear case where Bayesian principles should shine. Bayesian decision theory provides a framework for agentic systems that can help to maintain beliefs over task-relevant latent quantities, to update these beliefs from observed agentic and human-AI interactions, and to choose actions. Making LLMs themselves explicitly Bayesian belief-updating engines remains computationally intensive and conceptually nontrivial as a general modeling target. In contrast, this paper argues that coherent decision-making requires Bayesian principles at the orchestration level of the agentic system, not necessarily the LLM agent parameters. This paper articulates practical properties for Bayesian control that fit modern agentic AI systems and human-AI collaboration, and provides concrete examples and design patterns to illustrate how calibrated beliefs and utility-aware policies can improve agentic AI orchestration.
Theodore Papamarkou, Pierre Alquier, Matthias Bauer +27
Large language model (LLM) multi-agent systems typically rely on rigid orchestration, committing either to flat per-query routing or to hand-engineered task decomposition, so decomposition depth, worker choice, and inference budget are not jointly optimized under one objective. We introduce Uno-Orchestra, a unified orchestration policy that selectively decomposes a task and dispatches each subtask to an admissible (model, primitive) pair, with both decisions learned together from curated RL trajectories grounded in real worker interactions. Against 22 baselines on a 13-benchmark suite spanning math, code, knowledge, long-context, and agentic tool-use, Uno-Orchestra reaches 77.0% macro pass@1, roughly 16% above the strongest workflow baseline, at roughly an order of magnitude lower per-query cost, advancing the accuracy-efficiency frontier of selective delegation.