Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.
Figures & tables
Figure 1: MASkillBlender enables decentralized multi-humanoid whole-body coordination across diverse tasks and humanoid embodiments without requiring task-specific motion references.
Figure 2: Overview of MASkillBlender. (a) MASkillBlender enables diverse multi-humanoid coordination across different humanoid embodiments without task-specific motion references or extensive reward engineering. (b)-(c) For each task, relevant pre-trained primitive skills are selected and blended by a shared decentralized high-level policy. (d) The high-level policy is trained with MAPPO using permutation-based data augmentation.
Figure 3: Training curves of all methods on three Unitree H1 tasks.
Task
Metric
MASkillBlender
SkillBlender
TeamHOI
Carry
Return ( ↑ )
67.928 ± 1.563
65.855 ± 8.343
67.870 ± 9.935
Error box ( ↓ )
0.016 ± 0.008
0.025 ± 0.050
0.029 ± 0.102
Success ( ↑ )
100.0%
98.0%
96.0%
Push
Return ( ↑ )
61.137 ± 5.400
57.550 ± 5.011
57.991 ± 5.024
Error box ( ↓ )
0.116 ± 0.063
0.138 ± 0.069
0.136 ± 0.075
Success ( ↑ )
86.0%
84.0%
78.0%
Table 1: Quantitative comparison of MASkillBlender with SkillBlender and TeamHOI on the Unitree H1 tasks.
Figure 4: Qualitative comparison of MASkillBlender and TeamHOI on the Unitree H1 Carry task. Representative key frames from both methods are shown to illustrate the whole-body coordination behaviors during box lifting and transportation.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Number of Humanoids
Number of Reward Terms
Selected Primitive Skills
Walking
Reaching
Squatting
Carry
2
2
✓
✓
✓
Push
2
2
✓
✓
Move
3
3
✓
Appendix
Table 2: Task settings.
Task
Actor hidden dimensions
Critic hidden dimensions
Carry
[512, 256, 128]
[768, 256, 128]
Push
[512, 256, 128]
[1024, 512, 128]
Move
[512, 256, 128]
[768, 256, 128]
Appendix
Table 3: Network hyperparameters of MASkillBlender for three Unitree H1 tasks.
Task
Metric
MASkillBlender
MASkillBlender w/o Aug.
Carry
Return ( ↑ )
67.928 ± 1.563
45.791 ± 5.684
Error box ( ↓ )
0.016 ± 0.008
0.179 ± 0.068
Success ( ↑ )
100.0%
6.0%
Push
Return ( ↑ )
61.137 ± 5.400
57.130 ± 9.322
Error box ( ↓ )
0.116 ± 0.063
0.147 ± 0.098
Success ( ↑ )
86.0%
76.0%
Appendix
Table 4: Ablation study of permutation-based data augmentation on the Unitree H1 tasks.
Figure 5: Primitive skill ablation on the Unitree H1. Removing Squatting from Carry or Walking from Push substantially degrades learning performance.
Task
Metric
Nominal
Constrained
Carry
Return ( ↑ )
67.928 ± 1.563
63.462 ± 11.406
Error box ( ↓ )
0.016 ± 0.008
0.073 ± 0.201
Success ( ↑ )
100.0%
92.0%
Push
Return ( ↑ )
61.137 ± 5.400
52.744 ± 12.604
Error box ( ↓ )
0.116 ± 0.063
0.158 ± 0.089
Success ( ↑ )
86.0%
74.0%
Appendix
Table 5: Robustness evaluation of MASkillBlender under deployment noise and delays on the Unitree H1 tasks.
Figure 6: Long-horizon deployment from asymmetric initial configurations on the Unitree H1. The top row shows Carry and the bottom row shows Push . From left to right, the snapshots illustrate the asymmetric initial configuration, one humanoid arriving while the other is still approaching, both humanoids reaching the predefined task poses, and the final coordinated task execution.
Figure 7: Sim2Sim deployment of MASkillBlender in MuJoCo on (a) Carry and (b) Move .
Task
Metric
MuJoCo
Carry
Return ( ↑ )
66.787 ± 2.221
Error box ( ↓ )
0.026 ± 0.015
Move
Return ( ↑ )
86.323 ± 1.430
Error pos ( ↓ )
0.013 ± 0.005
Error head ( ↓ )
38.151 ± 30.906
Appendix
Table 6: Sim2Sim evaluation of MASkillBlender in MuJoCo on the Unitree H1 tasks.
Figure 8: Two-humanoid box-exchange task on the Unitree H1. The snapshots illustrate the execution sequence from the initial configuration to the final box exchange. From (a) to (f), the two humanoids first move their assigned boxes to the exchange region and then deliver the exchanged boxes to their final target positions.
Figure 9: Training curves of MASkillBlender on three Unitree G1 tasks.
Task
Metric
MASkillBlender
Carry
Return ( ↑ )
91.366 ± 5.339
Error box ( ↓ )
0.025 ± 0.043
Push
Return ( ↑ )
68.851 ± 13.408
Error box ( ↓ )
0.113 ± 0.108
Move
Return ( ↑ )
81.853 ± 2.115
Error pos ( ↓ )
0.039 ± 0.006
Appendix
Table 7: Performance of MASkillBlender on the Unitree G1 tasks.