For humanoids to be useful in everyday environments, they must perform a wide range of tasks that couple locomotion and manipulation. Existing approaches commonly acquire a loco-manipulation policy through reward engineering or demonstrations followed by task-specific training, making it costly to scale to new tasks. In this work, we propose a hierarchical approach to humanoid loco-manipulation that eliminates these per-task requirements. HuGo, Humanoid policy code Generation, uses a Large Language Model (LLM) to generate executable, closed-loop high-level policy code from a task description on top of a frozen low-level whole-body policy. Given the task, observation, and command specifications, the LLM constructs the task logic in code. HuGo then refines the policy from its rollouts using numerical trajectories and selected video frames to produce feedback and targeted code updates. Across five simulation tasks, using two different low-level policies, HuGo substantially outperforms a high-level reinforcement learning baseline and approaches the performance of a demonstration-based baseline. We achieve this level of performance without task-specific reward design or demonstration collection. We further demonstrate zero-shot transfer of simulation-generated policies to hardware and show that applying the same refinement loop to real-world rollouts can further improve transfer performance without expert demonstrations or policy retraining. Project website is https://iconlab.negarmehr.com/HuGo/
Figures & tables
Fig. 2 : The five simulated loco-manipulation tasks. All are situated in an indoor environment and randomize their initial conditions, so a policy must adapt its navigation path, approach pose, and command timing to each episode.
Low-level
Method
Low Gate Passing
Narrow Gate Passing
Push Button
Push Box
Lift Box
rl_isaac
High-level RL
11.33±1.70
0.0±0.0
7.33±1.70
1.67±1.25
0.0±0.0
HuGo (no refinement)
100.00±0.00
67.15±14.13
73.92±7.56
49.31±17.49
44.62±25.22
HuGo
100.00±0.00
96.20±1.31
79.97±1.68
89.33±11.37
75.92±16.44
sonic
HuGo (no refinement)
-
65.93±21.98
48.70±16.43
46.30±25.86
11.00±3.79
HuGo
-
97.54±1.97
85.78±7.89
61.79±12.18
24.91±6.12
TABLE I : Success rates (%) on the five-task suite with best-of- 10 selection. Reported success rates are evaluated on 100 held-out initial states not used for policy refinement or best-of-N selection. The same held-out initial states are used for all methods. HuGo (no refinement) selects among the same policies before refinement. sonic is not evaluated on LGP as it cannot change its height while walking.
PDH
CPB
HDMI [ 9 ]
99.95±2.21
98.4±12.6
HuGo
100.0±0.00
81.25±26.23
TABLE II : Success rates (%) on two HDMI tasks. HDMI trains a policy per task from a human demonstration, while our policies are generated from a task description without demonstrations. HuGo is reported with best-of- 10 selection.
Figure 4Figure 5Figure 6
Fig. 7 : Success rate at each refinement depth on hardware, starting from the policy that transfers worst. The simulation success rate of the same policy is shown for reference.
Fig. 8 : Example of one refinement depth on hardware, generated by the LLM without any human feedback. Top: executions at refinement depths 1 to 3 . Bottom: at depth 1 , the rollout feedback attributes the failure to the end-effector under-reaching the button, and the resulting code change introduces an explicit inward margin to compensate for the undershoot.
Department of Computer Science, University of Manchester, Manchester, UK · Human-Robot Interfaces and Interaction Laboratory, Italian Institute of Technology, Genoa, Italy