Humanoid Horizon: Extending Task Horizon in Whole-Body Loco-Manipulation via Parallel Training, Dynamic Starting, and Reward Gating
Organizations: The University of Manchester · University of Science and Technology of China · X-Humanoid · University of Warwick · Newcastle University
Abstract
Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy organizes scenes into concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero for the rest of the episode in later-stage streams when the immediately preceding object is displaced beyond a set threshold, so the shared policy learns not to disturb a just-placed object and earlier placements are preserved throughout the episode. Collectively, these strategies achieve per-stage success rates exceeding 80% on the two-object LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes). As the number of sequentially transported objects grows beyond two, success declines with the horizon, but the degradation is graceful relative to the sharp drop seen in all baselines.
Figures & tables
| Method | Succ1 | Succ2 | SuccAll | Dist 1 (m) | Dist 2 (m) |
|---|---|---|---|---|---|
| End-to-End RL | 2.3% | 0.0% | 0.0% | 2.05 | 1.97 |
| Curriculum RL | 88.3% | 48.5% | 47.2% | 0.26 | 1.01 |
| Hierarchical RL | 37.5% | 25.0% | 20.8% | 0.78 | 1.13 |
| HumanVLA Xu et al. (2024) | 42.3% | 36.7% | 29.9% | 0.65 | 0.97 |
| InterMimic Xu et al. (2025) | 20.6% | 9.8% | 7.5% | 1.32 | 1.82 |
| TokenHSI Pan et al. (2025) | 40.5% | 31.5% | 27.6% | 0.58 | 1.07 |
| Method | Succ1 | Succ2 | SuccAll | Dist 1 (m) | Dist 2 (m) |
|---|---|---|---|---|---|
| End-to-End RL | 1.4% | 0.0% | 0.0% | 2.37 | 2.28 |
| Curriculum RL | 75.5% | 40.4% | 39.5% | 0.36 | 1.11 |
| Hierarchical RL | 31.6% | 21.8% | 17.3% | 0.97 | 1.39 |
| HumanVLA Xu et al. (2024) | 36.3% | 31.3% | 25.1% | 0.72 | 1.16 |
| InterMimic Xu et al. (2025) | 17.2% | 8.2% | 5.9% | 1.47 | 2.03 |
| TokenHSI Pan et al. (2025) | 33.9% | 26.1% | 22.9% | 0.67 | 1.21 |
| Method | Succ1 | Succ2 | SuccAll | Dist 1 (m) | Dist 2 (m) |
|---|---|---|---|---|---|
| End-to-End RL | 1.3% | 0.0% | 0.0% | 2.50 | 2.49 |
| Curriculum RL | 71.5% | 41.1% | 40.7% | 0.37 | 1.18 |
| Hierarchical RL | 29.3% | 20.7% | 16.5% | 1.04 | 1.45 |
| HumanVLA Xu et al. (2024) | 34.2% | 29.8% | 23.5% | 0.70 | 1.26 |
| InterMimic Xu et al. (2025) | 16.3% | 8.4% | 6.2% | 1.64 | 2.22 |
| TokenHSI Pan et al. (2025) | 32.0% | 27.0% | 23.7% | 0.72 | 1.31 |
| Method | Succ1 | Succ2 | Succ3 | Succ4 | Succ5 | SuccAll |
|---|---|---|---|---|---|---|
| End-to-End RL | 1.5% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| Curriculum RL | 91.2% | 54.2% | 37.5% | 25.6% | 9.3% | 3.2% |
| Hierarchical RL | 38.6% | 27.3% | 16.6% | 3.8% | 0.0% | 0.0% |
| HumanVLA Xu et al. (2024) | 45.8% | 40.5% | 21.7% | 5.6% | 1.2% | 0.1% |
| InterMimic Xu et al. (2025) | 21.3% | 10.2% | 2.5% | 0.4% | 0.0% | 0.0% |
| TokenHSI Pan et al. (2025) | 41.7% | 32.7% | 18.9% | 4.9% | 1.1% | 0.1% |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | Value |
|---|---|
| PPO epochs per iteration | 4 |
| Horizon length (rollout steps) | 32 |
| Number of mini-batches | 8 |
| Discount factor | 0.98 |
| GAE | 0.95 |
| PPO clip | 0.2 |
| Loss term | Weight |
|---|---|
| Actor ( ) | 1.0 |
| Critic ( ) | 5.0 |
| Bound mu ( ) | 10.0 |
| Discriminator total ( ) | 5.0 |
| Disc prediction | 1.0 |
| Disc logit regularization | 0.01 |
| Reward term | Name | Formula | |
|---|---|---|---|
| 0.1 | Approach velocity | ||
| 0.1 | Approach position | ||
| 0.1 | Grasp proximity | ||
| 0.1 | Lift height | ||
| 0.2 | Object transport vel. | ||
| 0.2 | Object waypoint dist. |
| Hyperparameter | Value |
|---|---|
| Initial mixing coefficient | 0.998 |
| decay schedule | Exponential |
| Number of step iterations per epoch | 1 |
| Number of train iterations per epoch | 5 |
| Replay buffer size | 100,000 |
| Training batch size | 600 |
| Seed | success (SuccAll) | success_1 (Succ1) |
|---|---|---|
| Seed 0 | 87.7% | 92.2% |
| Seed 1 | 87.3% | 92.5% |
| Seed 2 | 87.8% | 92.8% |
| Mean Std | 87.6 0.3% | 92.5 0.3% |
| Succ1 | Succ2 | SuccAll | |
| Seen tasks (Table 1 ) | |||
| LHM-Humanoid | 1.55 | 1.76 | 1.91 |
| Humanoid Horizon | 0.94 | 0.13 | 0.94 |
| Unseen tasks (Table 2 ) | |||
| LHM-Humanoid | 3.27 | 3.78 | 3.78 |
| Humanoid Horizon | 0.71 | 1.43 | 1.24 |