Organizations: The University of Manchester · University of Science and Technology of China · X-Humanoid · University of Warwick · Newcastle University
Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy organizes N scenes into S concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero for the rest of the episode in later-stage streams when the immediately preceding object is displaced beyond a set threshold, so the shared policy learns not to disturb a just-placed object and earlier placements are preserved throughout the episode. Collectively, these strategies achieve per-stage success rates exceeding 80% on the two-object LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes). As the number of sequentially transported objects grows beyond two, success declines with the horizon, but the degradation is graceful relative to the sharp drop seen in all baselines.
Figures & tables
Figure 1: Humanoid Horizon uses a single policy for long-horizon whole-body loco-manipulation across cluttered bedroom, kitchen, living room, and warehouse scenes without resets. Each row shows one uninterrupted multi-object transport episode with sequential navigation, grasping, carrying, and placement.
Figure 2: Overview of the Humanoid Horizon training pipeline. (A) Parallel Training : S stage streams share one policy πθ and are optimized jointly, each initialized from its state buffer ( B1 – BS ). (B) Dynamic Starting : for each scene n , the terminal state from stage i overwrites the corresponding downstream buffer Bi+1n (one state per environment) to initialize stage i+1 in the next epoch; repeated overwrites across epochs broaden transition-state coverage over time. (C) Reward Gating : for stage- i streams ( i>1 ), reward is zeroed if object oi−1 moves beyond tolerance ϵ , preserving previously completed placements. (D) VLA Distillation : the RL teacher is distilled via DAgger into a vision-language-action student conditioned on egocentric RGB/depth observations and language instructions.
Method
Succ1 ↑
Succ2 ↑
SuccAll ↑
Dist 1 (m) ↓
Dist 2 (m) ↓
End-to-End RL
2.3%
0.0%
0.0%
2.05
1.97
Curriculum RL
88.3%
48.5%
47.2%
0.26
1.01
Hierarchical RL
37.5%
25.0%
20.8%
0.78
1.13
HumanVLA Xu et al. (2024)
42.3%
36.7%
29.9%
0.65
0.97
InterMimic Xu et al. (2025)
20.6%
9.8%
7.5%
1.32
1.82
TokenHSI Pan et al. (2025)
40.5%
31.5%
27.6%
0.58
1.07
Table 1: Results on 350 training tasks.
Method
Succ1 ↑
Succ2 ↑
SuccAll ↑
Dist 1 (m) ↓
Dist 2 (m) ↓
End-to-End RL
1.4%
0.0%
0.0%
2.37
2.28
Curriculum RL
75.5%
40.4%
39.5%
0.36
1.11
Hierarchical RL
31.6%
21.8%
17.3%
0.97
1.39
HumanVLA Xu et al. (2024)
36.3%
31.3%
25.1%
0.72
1.16
InterMimic Xu et al. (2025)
17.2%
8.2%
5.9%
1.47
2.03
TokenHSI Pan et al. (2025)
33.9%
26.1%
22.9%
0.67
1.21
Table 2: Results on 66 unseen tasks.
Method
Succ1 ↑
Succ2 ↑
SuccAll ↑
Dist 1 (m) ↓
Dist 2 (m) ↓
End-to-End RL
1.3%
0.0%
0.0%
2.50
2.49
Curriculum RL
71.5%
41.1%
40.7%
0.37
1.18
Hierarchical RL
29.3%
20.7%
16.5%
1.04
1.45
HumanVLA Xu et al. (2024)
34.2%
29.8%
23.5%
0.70
1.26
InterMimic Xu et al. (2025)
16.3%
8.4%
6.2%
1.64
2.22
TokenHSI Pan et al. (2025)
32.0%
27.0%
23.7%
0.72
1.31
Table 3: VLA extension results (all methods distilled via the same DAgger Ross et al. (2011) pipeline).
Method
Succ1 ↑
Succ2 ↑
Succ3 ↑
Succ4 ↑
Succ5 ↑
SuccAll ↑
End-to-End RL
1.5%
0.0%
0.0%
0.0%
0.0%
0.0%
Curriculum RL
91.2%
54.2%
37.5%
25.6%
9.3%
3.2%
Hierarchical RL
38.6%
27.3%
16.6%
3.8%
0.0%
0.0%
HumanVLA Xu et al. (2024)
45.8%
40.5%
21.7%
5.6%
1.2%
0.1%
InterMimic Xu et al. (2025)
21.3%
10.2%
2.5%
0.4%
0.0%
0.0%
TokenHSI Pan et al. (2025)
41.7%
32.7%
18.9%
4.9%
1.1%
0.1%
Table 4: More than two objects sequential transport results. LHM-Humanoid † denotes the variant trained with five teacher policies strictly following the original LHM-Humanoid Zhang et al. (2026) pipeline.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
PPO epochs per iteration
4
Horizon length (rollout steps)
32
Number of mini-batches
8
Discount factor γ
0.98
GAE λ
0.95
PPO clip ϵ
0.2
Appendix
Table 5: Teacher (RL) training hyperparameters.
Loss term
Weight
Actor ( wa )
1.0
Critic ( wc )
5.0
Bound mu ( wb )
10.0
Discriminator total ( wd )
5.0
Disc prediction
1.0
Disc logit regularization
0.01
Appendix
Table 6: Loss weights.
wk
Reward term
Name
Formula
0.1
Approach velocity
rvel
exp(−2(v∗−vproj)2)
0.1
Approach position
rpos
exp(−0.5⋅drobot→obj)
0.1
Grasp proximity
rgrasp
21∑hexp(−5⋅dh→obj)
0.1
Lift height
rheight
ztarget−zinitmin(zobj,ztarget)−zinit
0.2
Object transport vel.
robj_vel
exp(−2(vobj∗−vobj,proj)2)
0.2
Object waypoint dist.
robj_wp
exp(−1⋅dobj→waypoint)
Appendix
Table 7: Reward function components. All reward terms are bounded in [0,1] . The thresholds δr=0.5 m and δo=0.3 m gate the approach and grasp rewards once the robot/object is sufficiently close. When dobj→goal<δo , the approach, grasp, and height rewards are clamped to 1.0 to avoid interfering with placement.
Hyperparameter
Value
Initial mixing coefficient β0
0.998
β decay schedule
Exponential
Number of step iterations per epoch
1
Number of train iterations per epoch
5
Replay buffer size
100,000
Training batch size
600
Appendix
Table 8: DAgger distillation hyperparameters.
Figure 3: Training success curves across three random seeds (seed 0, 1, 2) over 16,000 epochs. Solid lines show per-seed success rates. All three seeds converge to similar final performance (SuccAll ≈ 87–89%), demonstrating that Humanoid Horizon is robust to random initialization. The initial rapid increase (epochs 0–2,000) corresponds to the pretrained single-object checkpoint’s adaptation to the parallel multi-object regime, followed by steady improvement through the interplay of dynamic starting and reward gating.
Seed
success (SuccAll)
success_1 (Succ1)
Seed 0
87.7%
92.2%
Seed 1
87.3%
92.5%
Seed 2
87.8%
92.8%
Mean ± Std
87.6 ± 0.3%
92.5 ± 0.3%
Appendix
Table 9: Final training metrics across three random seeds (last 100 epochs). success refers to the SuccAll rate logged during training, and success_1 refers to the Succ1 rate.
Succ1
Succ2
SuccAll
Seen tasks (Table 1 )
LHM-Humanoid
1.55
1.76
1.91
Humanoid Horizon
0.94
0.13
0.94
Unseen tasks (Table 2 )
LHM-Humanoid
3.27
3.78
3.78
Humanoid Horizon
0.71
1.43
1.24
Appendix
Table 10: Three-seed standard deviations of success rates, in percentage points.
Department of Computer Science, University of Manchester, Manchester, UK · Human-Robot Interfaces and Interaction Laboratory, Italian Institute of Technology, Genoa, Italy