Organizations: School of Computer Science, The University of Sydney, Australia. · College of Computer Science and Technology, Zhejiang University of Technology, China.
Collaborative manipulation requires robots to perform complementary actions as interactions unfold. We study single-policy decentralized collaboration: every robot runs the same policy from its visual observations and proprioception, without task prompts, identity labels, or inter-robot messages. The challenge is to learn complementary team behaviors within shared parameters and select appropriate actions from each robot's local observations. We introduce CoRE, which learns Collaboration-Role Experts from pooled multi-task, multi-robot demonstrations. Fused appearance and geometry provide local interaction evidence. Query-conditioned cross-attention experts provide adaptable prediction paths, which a local router combines at each action-chunk position. During training, an action-expert alignment loss supervises expert selection using relative forced-route prediction errors against demonstrations under fixed inputs, without role labels. Across simulation benchmarks, CoRE achieves the highest average performance among evaluated decentralized methods. Physical experiments demonstrate effective collaboration across diverse manipulation tasks and robustness to partner delays and slowdowns. Project page: https://aus.bot/research/core/.
Figures & tables
Fig. 1: Multi-robot collaboration paradigms. oti and ati are robot i ’s local observation and action at time t ; π and θ denote a policy and its parameters. (b) allows communication; (c) adds coordination inputs. The blue frame marks single-policy decentralized collaboration from local observations only.
Fig. 2: CoRE overview. A shared policy predicts action chunks from local observations. Collaboration-role experts augment shared cross-attention, with stepwise routing supervised by action–expert alignment. Shading illustrates schematic action error.
Fig. 3: Simulation benchmarks. Thirteen tasks from RoboFactory [ 3 ] , DuoBench [ 30 ] , and BiCoord [ 31 ] ; teams of two to four robots.
Fig. 4: Representative collaboration requirements illustrated through task sequences, progressing from left to right.
Method
Data
Models
Lift
Camera
3Stack
LPD
Photo
Avg.
RGB Regression
Task/robot
16
43 /100
19 /100
5 /100
0 /100
3 /100
14.0%
RGB Regression
Per task
5
86 /100
94 /100
26 /100
48 /100
14 /100
53.6%
RGB Regression
All tasks
1
87 /100
98 /100
3 /100
0 /100
14 /100
40.4%
CoRE (RGB)
All tasks
1
89 /100
100 /100
71 /100
86 /100
16 /100
72.4%
CoRE (Ours)
All tasks
1
99 /100
100 /100
99 /100
94 /100
29 /100
84.2%
TABLE II: Demonstration sharing. RoboFactory; metrics as in Table I . Data denotes pooling scope; Ours adds depth.
Fig. 5: Shared expert use during collaboration. (a) LPD: sequential object transfers, with robots numbered in transfer order. (b) Take Photo: asymmetric cooperation. Colors show action-history-weighted expert summaries; gray bars indicate grasp intervals. Snapshots align with dashed time guides.
Variant
Lift
Camera
3Stack
LPD
Photo
Avg.
RGB Regression
87 /100
98 /100
3 /100
0 /100
14 /100
40.4%
Stereo Regression
92 /100
100 /100
42 /100
20 /100
29 /100
56.6%
Stereo + MoE
100 /100
99 /100
68 /100
17 /100
28 /100
62.4%
Stereo + MoE + ARCA
96 /100
100 /100
81 /100
15 /100
29 /100
64.2%
CoRE (RGB)
89 /100
100 /100
71 /100
86 /100
16 /100
72.4%
CoRE (Ours)
99 /100
100 /100
99 /100
94 /100
29 /100
84.2%
TABLE III: Module ablations. RoboFactory with one shared policy. Task scores, averages, and ranking marks follow Table I .
E
K
Lift
Camera
3Stack
LPD
Photo
Avg.
2
2
96 /100
100 /100
99 /100
65 /100
41 /100
80.2%
4
1
87 /100
96 /100
82 /100
33 /100
15 /100
62.6%
4 *
2
99 /100
100 /100
99 /100
94 /100
29 /100
84.2%
4
4
100 /100
100 /100
59 /100
4 /100
31 /100
58.8%
8
2
99 /100
100 /100
98 /100
97 /100
23 /100
83.4%
TABLE IV: Expert count and routing sparsity. Settings and metrics as in Table III ; E experts, top- K routing; *: our default.
Fig. 6: Collaboration robustness. SR versus perturbed time (%) for random delay (Lift) and slowdown (ThreeStack); bands: 95% Wilson intervals.
Fig. 7: Real-world collaborative manipulation. Top: success rates over 30 trials per method and task; error bars denote 95% Wilson intervals. Bottom: representative task demonstrations.
Fig. 8: Physical collaboration. Coupled flipping followed by asymmetric joining in box assembly (top), and sequential block placement (bottom).
Fig. 9: Real-world collaboration robustness. Failure sequences under random partner pauses (Box, top) and 0.25× partner speed (Pan, bottom).
Fig. 10: Physical collaboration robustness. SR under periodic pauses (Box) and reduced speed (Pan); 30 trials per point, with 95% Wilson intervals.
Multi-robot collaboration allows robots to efficiently take on a wide range of tasks, from moving a couch through a doorway to assembling structures on a construction site. However, achieving such coordination in mobile multi-robot settings remains challenging: centralized methods conditioned on the combined observations of a team scale poorly with team size, and decentralized methods that train one policy per robot often require explicit alignment procedures or information sharing at inference time to overcome partial observability. Our key insight is that the visuomotor priors of pretrained vision-language-action (VLA) models should enable reactive, decentralized collaboration from each robot's local observations alone, without these inference-time assumptions. We propose CHORUS, a framework that adapts a single VLA backbone to control diverse, multi-robot teams. At inference time, each robot runs an independent copy of CHORUS, conditioned only on its own observations and a robot-identifying prompt. In real-world experiments including mobile tape measurement, library book handovers, and laundry basket lifting, CHORUS achieves a 64% point improvement over decentralized, from-scratch models, improves reactivity to teammate behavior by 40% points, and outperforms centralized baselines. Together, these results show that a shared VLA backbone is capable of achieving decentralized multi-robot collaboration, without per-robot policies or inter-robot communication at inference.
Multi-arm manipulation demands precise spatiotemporal coordination, yet many centralized approaches scale poorly as team size increases. To address this, we propose CLS-DP, a decentralized multi-agent framework that enables implicit coordination under partial observability without shared global views, explicit state information, or inter-agent communication. Under the centralized training and decentralized execution (CTDE) paradigm, CLS-DP distills privileged multi-agent dynamics into a latent space. At deployment, each agent infers a collaborative latent from its local RGB observation and a shared task instruction; it then conditions the diffusion denoising process on this latent. This design enables implicit coordination with a per-agent cost independent of team size. Across six RoboFactory benchmark tasks spanning two to four agents, CLS-DP achieves a 38% mean success rate, outperforming the best centralized baseline (20%) and a decentralized ablation without the collaborative latent (9%). It also maintains superior parameter efficiency across all agent configurations. Attribution maps show that an agent conditioned on the collaborative latent places high attribution on the joints and grippers of both itself and its teammates throughout execution. This suggests that the learned latent efficiently encodes collaborative dynamics from local observation, which facilitates implicit coordination in realistic settings characterized by partial observability.
Chanyoung Park, Minsung Yoon, Andrew Jeong +1
School of Computing at the Korea Advanced Institute of Science and Technology (KAIST), Daejeon, 34141, Republic of Korea
Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation policies directly in the physical world, avoiding the sim-to-real gap and enabling continuous policy refinement through human-in-the-loop interaction. Recent methods have demonstrated sample-efficient learning through human intervention but remain limited to small randomization ranges and encounter challenges with the non-stationarity induced by concurrently training multiple agents. To address these limitations, we introduce a unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA). This enables multiple actors to share a centralized multi-head critic. The critic is decomposed into task and grasp heads, corresponding to the sparse task reward and a potential-based grasping reward, respectively. We accordingly reformulate the critic and actor objectives to exploit the decomposed Q-values while explicitly accounting for the categorical action distribution of the discrete gripper policy. Experimental results demonstrate that the proposed framework substantially improves both sample efficiency and policy performance. We validate our approach on two robotic arms and a simulated humanoid robot across tennis ball and banana pick-and-place, pot reset, and simulated block relocation tasks under dimension-wise domain randomization, approximately 5-25x larger than those considered in prior work. Compared with a state-of-the-art baseline, our method improves the success rate from 60% to 80% on tennis ball pick-and-place, from 60% to 90% on banana pick-and-place, and from 25% to 95% on simulated block relocation, while also successfully accomplishing a task where the baseline consistently fails. Videos and more details are available at our project website: https://hil-harc.github.io/.
Changhao Li, Yifang Zhang, Heng Zhang +6
HHCM, Istituto Italiano di Tecnologia, Genoa, Italy · DIBRIS, University of Genova, Genova, Italy · Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genova, Italy +1