Pretrained robot policies offer strong manipulation skills but are typically limited to single-agent settings, where a robot acts in isolation. In this work, we study how to adapt pretrained single-agent diffusion policies to multi-agent settings using minimal collaborative data, co-optimizing for two key objectives: high coordination performance and single-agent skill retention. To this end, we introduce ALTER, an adaptation method for coordination on demand: the adapted policy coordinates with other robots when deployed in a team while remaining capable of acting independently when operating alone. Execution is decentralized: each robot acts only on its own visual observations, without explicit inter-agent communication. Our method trains a coordination head that predicts a residual denoiser to transform single-agent behavior into coordinated multi-agent behavior when necessary while also preserving single-agent capabilities. To preserve single-agent capabilities, we augment a small number of collaborative demonstrations with self-distilled data generated by the base policy during training of the residual denoiser. In simulation, ALTER achieves higher coordination success over our baselines while retaining much higher source-skill retention. In our hardware experiments, we find similar trends where ALTER better co-optimizes for coordination success and single-agent skill retention than the baselines.
Figures & tables
Fig. 3 : In the TwoArmPlaceWipe simulated task, one robot places and returns a tray while another wipes up dirt.
Multi-agent demos
Single-arm demos
Method
Checkpoint
Coordination success ↑
20
20
ALTER
500 k
35.0
20
20
FS
200 k
13.5
20
20
FT-mixed
200 k
33.0
20
0
FT-multi
200 k
29.5
40
40
ALTER
300 k
71.5
40
40
FS
900 k
28.0
TABLE I : TwoArmPlaceWipe coordination success (%, ↑ ) across demonstration budgets. Bold denotes the best result within each data regime.
Single-arm demos
Source success ↑
Multi-arm demos
Expert
Distilled rollouts
Method
Place- return
Wipe
Combined
0
400
0
Base policy
100.0
94.0
97.0
20
0
20
ALTER
99.0
95.0
97.0
20
0
20
FS
56.0
9.0
32.5
20
0
20
FT-mixed
55.0
37.0
46.0
20
0
0
FT-multi
0.0
40.0
20.0
TABLE II : Source-task success (%, ↑ ) after two-arm adaptation.
Size
Method
Total params. (M)
Trainable fraction (%)
Coordination success ↑
Source success ↑
XS
ALTER
17.881
2.4
25.0
97.0
XS
FS
17.881
100.0
12.0
34.5
S
ALTER
19.077
8.5
27.0
97.0
S
FS
19.076
100.0
10.0
37.5
M
ALTER
20.815
16.2
33.5
97.0
M
FS
20.815
100.0
10.5
39.5
TABLE III : Effect of coordination-head size on coordination and source-task success (%, ↑ ) in TwoArmPlaceWipe .
Fig. 4 : Hardware setup with two xArm7 robots, a lidded box, and a plush bird. One robot removes the lid, the other places the bird in the box, and the first robot replaces the lid.
Source success ↑
Method
Coordination success ↑
Lid removal
Lid replacement
Bird pick-place
Base policy
–
40/40
40/40
40/40
ALTER
18/20
19/20
20/20
20/20
FS
13/20
9/20
19/20
12/20
FT-mixed
18/20
11/20
19/20
17/20
TABLE IV : Hardware coordination and source success.
Robotics and Control Laboratory, School of Advanced Manufacturing and Robotics, and the State Key Laboratory of Turbulence and Complex Systems, Peking University, Beijing, 100871, China · National Innovation Institute of Defense Technology, Beijing 100071, China