Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.
Figures & tables
Figure 1: MASkillBlender enables decentralized multi-humanoid whole-body coordination across diverse tasks and humanoid embodiments without requiring task-specific motion references.
Figure 2: Overview of MASkillBlender. (a) MASkillBlender enables diverse multi-humanoid coordination across different humanoid embodiments without task-specific motion references or extensive reward engineering. (b)-(c) For each task, relevant pre-trained primitive skills are selected and blended by a shared decentralized high-level policy. (d) The high-level policy is trained with MAPPO using permutation-based data augmentation.
Figure 3: Training curves of all methods on three Unitree H1 tasks.
Task
Metric
MASkillBlender
SkillBlender
TeamHOI
Carry
Return ( ↑ )
67.928 ± 1.563
65.855 ± 8.343
67.870 ± 9.935
Error box ( ↓ )
0.016 ± 0.008
0.025 ± 0.050
0.029 ± 0.102
Success ( ↑ )
100.0%
98.0%
96.0%
Push
Return ( ↑ )
61.137 ± 5.400
57.550 ± 5.011
57.991 ± 5.024
Error box ( ↓ )
0.116 ± 0.063
0.138 ± 0.069
0.136 ± 0.075
Success ( ↑ )
86.0%
84.0%
78.0%
Table 1: Quantitative comparison of MASkillBlender with SkillBlender and TeamHOI on the Unitree H1 tasks.
Figure 4: Qualitative comparison of MASkillBlender and TeamHOI on the Unitree H1 Carry task. Representative key frames from both methods are shown to illustrate the whole-body coordination behaviors during box lifting and transportation.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Number of Humanoids
Number of Reward Terms
Selected Primitive Skills
Walking
Reaching
Squatting
Carry
2
2
✓
✓
✓
Push
2
2
✓
✓
Move
3
3
✓
Appendix
Table 2: Task settings.
Task
Actor hidden dimensions
Critic hidden dimensions
Carry
[512, 256, 128]
[768, 256, 128]
Push
[512, 256, 128]
[1024, 512, 128]
Move
[512, 256, 128]
[768, 256, 128]
Appendix
Table 3: Network hyperparameters of MASkillBlender for three Unitree H1 tasks.
Task
Metric
MASkillBlender
MASkillBlender w/o Aug.
Carry
Return ( ↑ )
67.928 ± 1.563
45.791 ± 5.684
Error box ( ↓ )
0.016 ± 0.008
0.179 ± 0.068
Success ( ↑ )
100.0%
6.0%
Push
Return ( ↑ )
61.137 ± 5.400
57.130 ± 9.322
Error box ( ↓ )
0.116 ± 0.063
0.147 ± 0.098
Success ( ↑ )
86.0%
76.0%
Appendix
Table 4: Ablation study of permutation-based data augmentation on the Unitree H1 tasks.
Figure 5: Primitive skill ablation on the Unitree H1. Removing Squatting from Carry or Walking from Push substantially degrades learning performance.
Task
Metric
Nominal
Constrained
Carry
Return ( ↑ )
67.928 ± 1.563
63.462 ± 11.406
Error box ( ↓ )
0.016 ± 0.008
0.073 ± 0.201
Success ( ↑ )
100.0%
92.0%
Push
Return ( ↑ )
61.137 ± 5.400
52.744 ± 12.604
Error box ( ↓ )
0.116 ± 0.063
0.158 ± 0.089
Success ( ↑ )
86.0%
74.0%
Appendix
Table 5: Robustness evaluation of MASkillBlender under deployment noise and delays on the Unitree H1 tasks.
Figure 6: Long-horizon deployment from asymmetric initial configurations on the Unitree H1. The top row shows Carry and the bottom row shows Push . From left to right, the snapshots illustrate the asymmetric initial configuration, one humanoid arriving while the other is still approaching, both humanoids reaching the predefined task poses, and the final coordinated task execution.
Figure 7: Sim2Sim deployment of MASkillBlender in MuJoCo on (a) Carry and (b) Move .
Task
Metric
MuJoCo
Carry
Return ( ↑ )
66.787 ± 2.221
Error box ( ↓ )
0.026 ± 0.015
Move
Return ( ↑ )
86.323 ± 1.430
Error pos ( ↓ )
0.013 ± 0.005
Error head ( ↓ )
38.151 ± 30.906
Appendix
Table 6: Sim2Sim evaluation of MASkillBlender in MuJoCo on the Unitree H1 tasks.
Figure 8: Two-humanoid box-exchange task on the Unitree H1. The snapshots illustrate the execution sequence from the initial configuration to the final box exchange. From (a) to (f), the two humanoids first move their assigned boxes to the exchange region and then deliver the exchanged boxes to their final target positions.
Figure 9: Training curves of MASkillBlender on three Unitree G1 tasks.
Task
Metric
MASkillBlender
Carry
Return ( ↑ )
91.366 ± 5.339
Error box ( ↓ )
0.025 ± 0.043
Push
Return ( ↑ )
68.851 ± 13.408
Error box ( ↓ )
0.113 ± 0.108
Move
Return ( ↑ )
81.853 ± 2.115
Error pos ( ↓ )
0.039 ± 0.006
Appendix
Table 7: Performance of MASkillBlender on the Unitree G1 tasks.
Humanoid soccer is a challenging testbed for dynamic whole-body control, requiring robots to coordinate balance, locomotion, object interaction, and skill switching over long horizons. Existing humanoid sports methods often rely on task-specific multi-stage pipelines, making it difficult to jointly learn and compose multiple object-interactive skills within a single deployable policy. To address this, we present SkillX, a unified reinforcement learning framework that learns and composes multiple atomic soccer skills through a single command-conditioned policy. SkillX integrates three core designs: skill-specific adversarial motion priors, skill-specific critics, and an object-aware temporal encoder, enabling the robot to execute atomic skills and transition among them such as dribbling, trapping, and shooting. Experiments in simulation and on a real Noetix E1 humanoid demonstrate robust multi-skill execution, long-horizon skill composition, and successful sim-to-real deployment.
Multi-arm manipulation demands precise spatiotemporal coordination, yet many centralized approaches scale poorly as team size increases. To address this, we propose CLS-DP, a decentralized multi-agent framework that enables implicit coordination under partial observability without shared global views, explicit state information, or inter-agent communication. Under the centralized training and decentralized execution (CTDE) paradigm, CLS-DP distills privileged multi-agent dynamics into a latent space. At deployment, each agent infers a collaborative latent from its local RGB observation and a shared task instruction; it then conditions the diffusion denoising process on this latent. This design enables implicit coordination with a per-agent cost independent of team size. Across six RoboFactory benchmark tasks spanning two to four agents, CLS-DP achieves a 38% mean success rate, outperforming the best centralized baseline (20%) and a decentralized ablation without the collaborative latent (9%). It also maintains superior parameter efficiency across all agent configurations. Attribution maps show that an agent conditioned on the collaborative latent places high attribution on the joints and grippers of both itself and its teammates throughout execution. This suggests that the learned latent efficiently encodes collaborative dynamics from local observation, which facilitates implicit coordination in realistic settings characterized by partial observability.
Chanyoung Park, Minsung Yoon, Andrew Jeong +1
School of Computing at the Korea Advanced Institute of Science and Technology (KAIST), Daejeon, 34141, Republic of Korea
We study cooperative multi-humanoid pickup and transport of objects with varying size, weight, and geometry, requiring robot teams of different sizes. Our approach uses decentralized object-centric control, where each humanoid is assigned a local attachment region on the shared object and learns to realize pickup and transport through gripperless bimanual pinching. This attachment-based interface provides a common control abstraction spanning single-robot pickup, cooperative multi-robot transport, and robot-to-robot handover, without per-task redesign. We find that policies trained only on single-robot pickup already transfer nontrivially to cooperative settings, suggesting that this abstraction captures much of the structure needed for coordination. At the same time, explicit multi-robot training further improves performance, showing that shared-object coupling introduces coordination dynamics that are beneficial to learn directly. We validate the approach in simulation across varying team sizes and object geometries, and demonstrate sim-to-real transfer on hardware, where the learned controllers enable real humanoids to perform cooperative manipulation tasks.
Bikram Pandit, Mohitvishnu S. Gadde, Aayam Kumar Shrestha +1
Dynamics Robotics and AI Lab (DRAIL), Oregon State University, Corvallis, OR, USA