One for All, All for One: Coordinated Multi-Agent Diffusion Steering via Stochastic Optimal Control
Organizations: UCL · TUM & MCML · CUHK · KCL · DESY · TUM & MCML & CIFAR · Proxima Bio · University of Edinburgh & CIFAR
Abstract
Deep generative models often produce structured outputs composed of interacting components. Modelling these outputs with a single model requires learning both the component distributions and their interactions. We pursue a modular alternative: reuse independently trained component generators and learn only how to coordinate them to produce coherent structured outputs. Our framework, Coordinated Multi-Agent Diffusion Steering (CMDS), treats frozen pretrained diffusion models as reusable generative primitives and coordinates their reverse processes through a learned control. We formulate coordination as a stochastic optimal control problem, balancing an assembly-level reward that specifies the desired properties of the combined output against deviations from the pretrained dynamics. The learned control amortises this optimisation, allowing reuse across new task instances. Experiments show that CMDS can recover a known target distribution, satisfy different spatial constraints with the same trained control, and recover individual sources from degraded mixtures. Across multi-agent maze navigation, articulated robot planning, and text-conditioned human motion, CMDS turns frozen models into coordinated multi-agent generators.
Figures & tables
| Medium | Large | |||||||
| Method | Pool 1 | Pool 2 | Pool 1 | Pool 2 | Pool 1 | Pool 2 | Pool 1 | Pool 2 |
| Reference | 0.174 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} | 0.195 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.008} | 0.068 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} | 0.010 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} | 0.195 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.016} | 0.219 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.000} | 0.039 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.000} | 0.049 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} |
| Best-of-256 | 0.216 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} | 0.263 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.016} | 0.076 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} | 0.044 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} | 0.232 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.009} | 0.328 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.021} | 0.076 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} | 0.070 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.000} |
| FK-Steering | 0.797 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.028} | 0.820 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.034} | 0.398 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.036} | 0.331 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.016} | 0.924 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.030} | 0.953 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.008} | 0.622 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.009} | 0.729 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.023} |
| DPS | 0.964 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.024} | 0.956 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.016} | 0.958 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.025} | 0.919 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.033} | 0.984 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.016} | 0.977 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.000} | 0.880 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.030} | 0.927 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.009} |
| Method | Strict success (%, ) | Planning time (s, ) |
|---|---|---|
| PBDM | 60.9 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 8.0} | 0.427 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.003} |
| Dual KUKA + guidance | 74.3 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 5.7} | 0.356 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.002} |
| Dual KUKA prior | 36.3 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 8.7} | 0.274 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.033} |
| Single-arm priors + guidance | 51.0 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.6} | 0.355 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.001} |
| Reference single-arm priors | 11.4 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 1.5} | 0.193 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.002} |
| CMDS +TRG | 77.3 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.7} | 0.368 \,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.002} |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Notation |
|---|---|
| Variables and processes | |
| Abstract component & assembly | |
| (Uncontrolled) reference reverse process | |
| Generic controlled reverse process | |
| Parametrised controlled reverse process | |
| Stop-gradient controlled rollout | |
| Method | Optim. scheme | ELBO | Valid | |
|---|---|---|---|---|
| (1-square) | Adj. Match. | -1303.1 | -7.2 | 0.0% |
| Disc. Adj. | -1309.5 | -7.2 | 0.0% | |
| (2-square) | Adj. Match. | -289.1 | -0.7 | 18.9% |
| Disc. Adj. | -173.4 | -0.7 | 25.4% | |
| CMDS | Adj. Match. | -74.3 | -1.6 | 99.2% |
| Disc. Adj. | -61.2 | -1.6 | 98.5% |
| Method | Optim. scheme | ELBO | Valid | |
|---|---|---|---|---|
| (1-square) | Adj. Match. | -1706.8 | -22.8 | 0.0% |
| Disc. Adj. | -1565.0 | -22.8 | 0.0% | |
| (2-square) | Adj. Match. | -1561.7 | -4.5 | 8.3% |
| Disc. Adj. | -381.4 | -4.5 | 9.6% | |
| CMDS | Adj. Match. | -156.7 | -5.4 | 97.9% |
| Disc. Adj. | -190.9 | -5.4 | 96.8% |
| Training scheme | ELBO | Valid | |||||
|---|---|---|---|---|---|---|---|
| Avg. | |||||||
| Adj. Match. | -172.2 | -5.4 | 95.8% | 95.4% | 97.3% | 97.1% | 96.4% |
| Disc. Adj. | -158.5 | -5.4 | 95.2% | 95.7% | 97.2% | 95.5% | 95.9% |
| Medium ( ) | Medium ( ) | Large ( ) | Large ( ) | |||||
|---|---|---|---|---|---|---|---|---|
| Matching loss | Pool 1 | Pool 2 | Pool 1 | Pool 2 | Pool 1 | Pool 2 | Pool 1 | Pool 2 |
| Squared | 0.940\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.020} | 0.884\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.048} | \mathbf{0.806}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.155} | 0.667\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.151} | 0.833\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.041} | 0.792\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.042} | 0.723\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.157} | 0.606\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.152} |
| Pseudo-Huber | \mathbf{0.988}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.012} | \mathbf{0.933}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.032} | 0.805\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.098} | 0.665\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.169} | \mathbf{0.907}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.037} | \mathbf{0.833}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.045} | \mathbf{0.873}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.027} | |
| Medium ( ) | Medium ( ) | Large ( ) | Large ( ) | |||||
| Method | Pool 1 | Pool 2 | Pool 1 | Pool 2 | Pool 1 | Pool 2 | Pool 1 | Pool 2 |
| CMDS +TRG | \mathbf{0.998}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.003} | \mathbf{0.983}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.015} | \mathbf{0.984}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} | \mathbf{0.947}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.006} | \mathbf{0.973}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.009} | \mathbf{0.906}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.015} | \mathbf{0.975}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.006} | \mathbf{0.969}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.005} |
| Population | Construction | Evaluation tasks |
|---|---|---|
| In-domain | Released dual-arm scene and endpoints | 413 |
| Shifted obstacles | Horizontal Gaussian shifts, standard deviation 0.10 m | 191 |
| Larger obstacles | Each box half-extent multiplied by 1.2 | 121 |
| Wider bases | Base positions moved to m | 208 |
| Three arms | Third base at m, identity rotation | 172 |
| Setting | Final configuration |
|---|---|
| Trajectory / reverse steps | 56 knots / 10 DDIM steps |
| Stochasticity / prior conditioning | / |
| Controller | Width 344, 4 heads, 2 axial blocks |
| Feed-forward / pair-bias hidden width | controller width / 32 |
| Optimiser | Adam, , |
| Updates / batch / microbatch | 20,000 / 256 joint queries / 128 |
| Primitive | Joint endpoints | Radius (m) |
|---|---|---|
| Pelvis | 0.120 | |
| Torso | 0.150 | |
| Shoulder span | 0.100 | |
| Upper torso–head | 0.100 | |
| Head | 0.110 | |
| Upper arms | 0.055 |
| Setting | Recorded value |
|---|---|
| , training | |
| , baselines | |
| Outer SOC reward multiplier | , control training only |
| Optimiser / updates / learning rate | Adam / / constant |
| Method | Joint success (%) | Collision-free (%) | Peak body overlap (cm) | All-person goals (%) | Goal error (m) | Kinematic validity (%) | Time (s/assembly) |
|---|---|---|---|---|---|---|---|
| Best-of-256 | 0.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 16.7\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 5.9} | 8.860\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 1.086} | 0.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 2.456\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.146} | 83.3\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 6.6} | 69.895\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.292} |
| DPS | 0.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 61.7\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 10.8} | 3.335\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 1.049} | 0.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 3.751\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.284} | 57.5\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 6.8} | 8.566\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.018} |
| PCD | 4.2\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 2.9} | 8.3\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 2.9} | 19.002\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 1.865} | \mathbf{100.0}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 0.500\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.000} | 27.5\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 4.8} | 9.985\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.064} |
| PCD++ | 75.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 6.6} | \mathbf{98.3}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 2.3} | \mathbf{0.000}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.000} | 98.3\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 2.3} | 0.487\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.001} | 75.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 6.6} | 145.335\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 4.460} |
| CMDS | \mathbf{90.0}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 6.3} | 90.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 6.3} | 0.740\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.351} | \mathbf{100.0}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | \mathbf{0.158}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.008} | \mathbf{100.0}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | \mathbf{0.541}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.002} |