Organizations: School of Computer Science, The University of Sydney, Australia. · College of Computer Science and Technology, Zhejiang University of Technology, China.
Collaborative manipulation requires robots to perform complementary actions as interactions unfold. We study single-policy decentralized collaboration: every robot runs the same policy from its visual observations and proprioception, without task prompts, identity labels, or inter-robot messages. The challenge is to learn complementary team behaviors within shared parameters and select appropriate actions from each robot's local observations. We introduce CoRE, which learns Collaboration-Role Experts from pooled multi-task, multi-robot demonstrations. Fused appearance and geometry provide local interaction evidence. Query-conditioned cross-attention experts provide adaptable prediction paths, which a local router combines at each action-chunk position. During training, an action-expert alignment loss supervises expert selection using relative forced-route prediction errors against demonstrations under fixed inputs, without role labels. Across simulation benchmarks, CoRE achieves the highest average performance among evaluated decentralized methods. Physical experiments demonstrate effective collaboration across diverse manipulation tasks and robustness to partner delays and slowdowns. Project page: https://aus.bot/research/core/.
Figures & tables
Fig. 1: Multi-robot collaboration paradigms. oti and ati are robot i ’s local observation and action at time t ; π and θ denote a policy and its parameters. (b) allows communication; (c) adds coordination inputs. The blue frame marks single-policy decentralized collaboration from local observations only.
Fig. 2: CoRE overview. A shared policy predicts action chunks from local observations. Collaboration-role experts augment shared cross-attention, with stepwise routing supervised by action–expert alignment. Shading illustrates schematic action error.
Fig. 3: Simulation benchmarks. Thirteen tasks from RoboFactory [ 3 ] , DuoBench [ 30 ] , and BiCoord [ 31 ] ; teams of two to four robots.
Fig. 4: Representative collaboration requirements illustrated through task sequences, progressing from left to right.
Method
Data
Models
Lift
Camera
3Stack
LPD
Photo
Avg.
RGB Regression
Task/robot
16
43 /100
19 /100
5 /100
0 /100
3 /100
14.0%
RGB Regression
Per task
5
86 /100
94 /100
26 /100
48 /100
14 /100
53.6%
RGB Regression
All tasks
1
87 /100
98 /100
3 /100
0 /100
14 /100
40.4%
CoRE (RGB)
All tasks
1
89 /100
100 /100
71 /100
86 /100
16 /100
72.4%
CoRE (Ours)
All tasks
1
99 /100
100 /100
99 /100
94 /100
29 /100
84.2%
TABLE II: Demonstration sharing. RoboFactory; metrics as in Table I . Data denotes pooling scope; Ours adds depth.
Fig. 5: Shared expert use during collaboration. (a) LPD: sequential object transfers, with robots numbered in transfer order. (b) Take Photo: asymmetric cooperation. Colors show action-history-weighted expert summaries; gray bars indicate grasp intervals. Snapshots align with dashed time guides.
Variant
Lift
Camera
3Stack
LPD
Photo
Avg.
RGB Regression
87 /100
98 /100
3 /100
0 /100
14 /100
40.4%
Stereo Regression
92 /100
100 /100
42 /100
20 /100
29 /100
56.6%
Stereo + MoE
100 /100
99 /100
68 /100
17 /100
28 /100
62.4%
Stereo + MoE + ARCA
96 /100
100 /100
81 /100
15 /100
29 /100
64.2%
CoRE (RGB)
89 /100
100 /100
71 /100
86 /100
16 /100
72.4%
CoRE (Ours)
99 /100
100 /100
99 /100
94 /100
29 /100
84.2%
TABLE III: Module ablations. RoboFactory with one shared policy. Task scores, averages, and ranking marks follow Table I .
E
K
Lift
Camera
3Stack
LPD
Photo
Avg.
2
2
96 /100
100 /100
99 /100
65 /100
41 /100
80.2%
4
1
87 /100
96 /100
82 /100
33 /100
15 /100
62.6%
4 *
2
99 /100
100 /100
99 /100
94 /100
29 /100
84.2%
4
4
100 /100
100 /100
59 /100
4 /100
31 /100
58.8%
8
2
99 /100
100 /100
98 /100
97 /100
23 /100
83.4%
TABLE IV: Expert count and routing sparsity. Settings and metrics as in Table III ; E experts, top- K routing; *: our default.
Fig. 6: Collaboration robustness. SR versus perturbed time (%) for random delay (Lift) and slowdown (ThreeStack); bands: 95% Wilson intervals.
Fig. 7: Real-world collaborative manipulation. Top: success rates over 30 trials per method and task; error bars denote 95% Wilson intervals. Bottom: representative task demonstrations.
Fig. 8: Physical collaboration. Coupled flipping followed by asymmetric joining in box assembly (top), and sequential block placement (bottom).
Fig. 9: Real-world collaboration robustness. Failure sequences under random partner pauses (Box, top) and 0.25× partner speed (Pan, bottom).
Fig. 10: Physical collaboration robustness. SR under periodic pauses (Box) and reduced speed (Pan); 30 trials per point, with 95% Wilson intervals.
HHCM, Istituto Italiano di Tecnologia, Genoa, Italy · DIBRIS, University of Genova, Genova, Italy · Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genova, Italy +1