We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.
Figures & tables
Figure 1: InterEvolve adapts a frozen whole-body controller with novel use. The controller trained on human-object interaction (HOI), such as lifting and pushing a box, is asked to rotate a box around in place, a behavior absent from its training data (Sec. D.7 and Figure 11 ). A human-designed reward tries to hack this behavior but fails. The reward program that InterEvolve evolves compensate with novel body used to succeed. Bottom: InterEvolve can achieve versatile evolved skills in simulation and on a real robot. More demos are on the project page .
Figure 2: InterEvolve overview. (a) Given a task and scene, an LLM agent writes a staged reward program, drawing on verified programs in the skill library. (b) The forward-backward model scores the active stage’s reward over a fixed state bank and projects it into a latent prompt, which a frozen body prior and a pretrained object residual execute without retraining. (c) A fixed verifier scores parallel rollouts, and criterion-level feedback drives the next structural revision (outer loop), while CMA-ES tunes the program constants (inner loop).
Table 3
Figure 3: Dexterous hands and long-horizon composition. (a) The G1 with Inspire hands, tracking a reference with the tracking reward calibrated by InterEvolve. Left : whole-body view at the moment of lift; Right : close-ups of the manipulation phases. (b) Arranging six boxes by the evolution. Left : top-down trace of the robot (gray) and each box from its start (outlined square) to its target cell (dashed circle); numbers give the push order. Right : the robot pushing each box.
Search components
Results
CMA-ES tuning
Targeted edits
Multi- scenario
Scene context
Multi- stage
SR (%) ↑
Earned (%) ↑
GPU-h ↓
Tokens (M) ↓
✓
✓
✓
✓
✓
86.5
95.6
2.1
0.23
✗
✓
✓
✓
✓
51.6
89.9
1.7
0.27
✓
✗
✓
✓
✓
68.9
92.3
2.7
0.20
✓
✓
✗
✓
✓
68.4
89.9
1.0
0.17
✓
✓
✓
✗
✓
78.7
93.2
2.0
0.21
Table 3: Ablation of the search design. Stages and numerical calibration are most important designs. Each row disables one component (✗) of the full system; definitions are in Sec. D.5 .
Figure 4: Autonomous execution of real-world G1 with onboard camera and task intent of kicking a box (top) and pushing a box (bottom). Each row runs a program evolved in simulation. The first four columns show the run from a third-person view, left to right in time. The last two columns are the robot’s onboard color and depth images as it approaches the box, overlaid with the FoundationPose estimate (yellow) and the filtered pose that the controller receives (green).
Table 7
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Interaction encoding
Dim.
Eh↓
Eo↓
SR ↑
Δ SR
Object state + link geometry
105
27.20
30.80
60
n/a
No object state
90
28.57
49.20
20
−40
No link geometry
15
25.05
60.83
12
−48
Appendix
Table 6: Interaction observations. Ablations remove the object state or the link geometry. Δ SR is relative to the full 105-dimensional encoding, which is the main model of Table 2 (without replanning).
Model
Trainable parameters
Eh↓
Eo↓
SR ↑
From scratch
843 M
60.96
78.21
2
Residual on the body prior
MLP 256×2
14.1 M
36.19
50.81
28
MLP 512×2
30.5 M
31.51
42.58
30
MLP 1024×2
79.5 M
29.22
37.42
46
Full-width residual
852.0 M
27.20
30.80
60
Appendix
Table 7: Residual capacity. MLP width × depth of the residual branch; the full-width residual mirrors the prior’s module architectures and is the main model of Table 2 (without replanning). Parameter counts include all trainable modules. Errors and SR follow Table 2 .
Anchor
Addition
kp
SR (%) ↑
Eh↓
Eo↓
None
–
–
60
27.20
30.80
Anchor
Root
–
2
68
19.56
25.27
Object
–
3
72
19.80
24.92
Root and object
–
2, 3
68
18.24
24.10
Correction gain kp
Appendix
Table 8: Drift-correction ablation. Tracking protocol of Table 2 ; Eh and Eo in cm. The anchor is the quantity steered toward the reference, kp is the correction gain, and the shaded row is the configuration used in Table 2 .
Figure 5: World-frame drift over time. Horizontal distance between the simulated and reference root (left) and object (right) for native tracking and the object-anchored drift correction, on the 50 largebox clips of Table 2 . Lines show the median and bands the 25–75% range over all 50 motions; each motion is counted to the end of its clip, and after its clip ends its final value is held. Without correction the median root drift grows to 0.29 m at 4 s and 0.32 m by the end, against 0.16 m and 0.18 m with the correction, which levels off after about 4 s. The two curves coincide for the first second because the correction has a 0.10 m deadband.
β
State count
SR (%) ↑
N90
Projection latency (ms) ↓
Peak memory (MiB) ↓
Reward tilt
0 (uniform)
50k
0.0
45,000
–
–
1
50k
15.6
28,674
–
–
3
50k
35.9
10,655
–
–
10
50k
59.4
2,077
0.92
63
30
50k
28.1
3
–
–
Appendix
Table 9: Reward-inference bank diagnostics on the feet-only task. SR is the task success rate of the fixed program; N90 is the number of bank states that carry 90% of the weight at a frame, averaged over the frames of a rollout and reported as the median over the 64 test scenarios. Latency is per reward projection; peak memory covers resident bank features and projection workspace.
Figure 6: Reward tilt, bank size, and cost on the feet-only task. (a) SR against the reward tilt β with 50k states. (b) SR against bank size at β=10 . (c) Peak memory (solid) and projection latency (dashed) against bank size.
Component
Symbol
Role
Example edit
Reward code
rj
Desired body-object interaction
Upward → horizontal object motion
Constants
θ
Scale and target of an objective
Target height, contact offset
Stages
j
Intermediate objectives
Split acquisition from transport
Completion
gj
Completing or switching a stage
Require support before transport
Stage memory
–
Context captured at runtime
Displacement since stage entry
Appendix
Table 10: Editable components of a reward program. Each edit is judged by its executed outcome.
Figure 7: Reward-program evolution on the kick task. Each row shows the current best program after the initial proposal, rounds 1–3, and the final round (top to bottom), the kick programs evaluated in Table 5 ; frames run left to right. Revised objectives change the contact strategy across rounds.
Figure 8: The eight basic skills. Final evolved program of each task family, frames left to right. Whole-body motion links object acquisition, transport, and placement across carrying, lifting, tipping, kicking, and obstacle-aware pushing.
Figure 9: Composite tasks. The three composite tasks of Table 5 : relocating a box between supports, stacking one box on another, and carrying, placing, then kicking a box. Frames run left to right.
ULTRA
BFM-Zero
InterEvolve
+ Evolving
SR (%) ↑
66
8
60
72
Eh (cm) ↓
15.68
28.36
27.20
19.80
Eo (cm) ↓
21.45
62.33
30.80
24.92
Orientation error ( ∘ ) ↓
37.8
47.6
23.4
26.2
Contact retention (%) ↑
92
2
89
89
Slip, sim / ref. (cm/s)
27.7 / 19.4
20.0 / 26.4
20.0 / 20.8
19.6 / 20.5
Appendix
Table 11: Supplementary tracking metrics on the clips of Table 2 . Every metric is averaged over three evaluation runs of the 50 clips, with rates over clips rounded to integers, and the rows below the first block come from recordings in which every clip runs to its end without early termination. Orientation error : object orientation error, averaged over the whole clip. Contact retention : fraction of reference-contact frames in which the simulated hand is also in contact, with contact meaning a hand vertex within 3 cm of the object surface. Slip : wrist speed in the object frame during contact, in simulation and on the reference. Penetration , object below ground : frames deeper than 2 cm. Falls : pelvis below 0.35 m or tilted beyond 60∘ .
Task
Interaction
Key constraint
Push to mark
Push object to a ground target
Keep it on the floor, stop near target
Carry at mid height
Acquire, raise, and carry
Hold a waist-relative height band
Carry at chest height
Acquire, raise, and carry
Hold a body-relative height
Tip onto a new face
Turn it over about a horizontal axis
Large tilt, left standing on the new face
Push through a gate
Push between two immovable walls
Hold a straight heading through the gap
Push around an obstacle
Push past a pillar in the way
Deviate, then recover the line
Appendix
Table 12: Task families for program search. The verifier scores outcomes and constraints independently of the agent-written reward.
Task family
Human + CMA
Initial agent + CMA
InterEvolve
Push to mark
96.9
100.0
100.0
Carry at mid height
0.0
40.6
100.0
Carry at chest height
0.0
0.0
100.0
Tip onto a new face
0.0
0.0
95.3
Push through a gate
0.0
25.0
73.4
Push around an obstacle
0.0
4.7
71.9
Appendix
Table 13: Success per task family (%). CMA denotes CMA-ES. The macro-average weights families equally.
Figure 10: Push to mark under new goals and starts , with the selected program unchanged. (a) From one start, targets at bearings from −45∘ to +45∘ , 2.5 m away; rings mark the targets. (b) The box starts 0.55, 1.2, or 1.8 m from the robot at bearings of −40∘ , 0∘ , and +40∘ , and is pushed to one target (red ring). Lines trace the box; each rollout shown is one of 16 per setting.
Task family
Large box
Plastic box
Small box
Suitcase
Push to mark
100.0
0.0 (4.7) → 98.4
0.0 (0.0) → 98.4
0.0 (25.0) → 92.2
Carry at mid height
100.0
15.6 → 84.4
35.9 → 96.9
76.6 → 90.6
Carry at chest height
100.0
18.8 → 98.4
29.7 → 95.3
98.4 †
Tip onto a new face
95.3
78.1 → 93.8
28.1 → 89.1
31.2 → 96.9
Push through a gate
73.4
40.6 → 98.4
56.2 → 85.9
68.8 → 90.6
Push around an obstacle
71.9
32.8 †
1.6 → 84.4
96.9 †
Appendix
Table 14: Selected programs on other boxes. Each program of Table 13 is searched on the large box and run on the three other box-like OMOMO objects the controller was trained on, first unchanged (left of the arrow) and then adapted to each box by hand; † marks boxes where adaptation did not help and the unchanged program is kept. SR (%) over 64 test scenarios, as in Table 13 ; the large-box column is a rerun of the same protocol and reproduces that table exactly.
Task
Scene and interaction
Verifier criteria
Relocate between supports
The box starts on a 0.3 m support in front of the robot. The robot lifts it, turns around, carries it 2.9 m to a 0.6 m support behind it, sets it down, and lets go.
Stays upright; holds the box for a large part of the episode; lifts it above the destination height; ends within 0.3 m of the mark on the support; hands clear of the box at the end.
Stack two boxes
Box A is on the floor 0.5 m ahead; box B, of the same size and free to move, is 1.9 m to the left. The robot lifts A, carries it above B, lowers it onto B’s top face, and lets go.
Stays upright; lifts A clear of B’s top; sustained two-handed contact; A stays level; ends within 0.25 m of B’s top (the target follows B if B is pushed); hands clear; A at rest.
Carry, place, and kick
The box is on the floor 0.5 m ahead. The robot pushes it about 1.0 m along the ground with its hands, lets go, and kicks it with its feet a further 1.5 m to a mark 3.0 m ahead.
Stays upright; box stays low (no lift); sustained hand contact during the push; at least one foot contact; box travels at least 1.0 m after the last hand contact; hands clear; ends within 0.3 m of the mark.
Appendix
Table 15: Composite tasks for library reuse. Distances are measured from the box’s starting position.
Stored program, no search
SR (%)
Kicking to mark
6.3
Push to mark
0.0
Carry at mid height
0.0
Carry at chest height
0.0
Push around an obstacle
0.0
Push through a gate
0.0
Appendix
Table 16: Library ablations on carry, place, and kick , extending Table 5 . Every row is scored on the same 64 test scenarios, none of which any search used; SR is the percentage of scenarios solved, where a scenario is solved when one of three rollouts meets every criterion. Left: each stored program run as written, without search. Right: three search rounds from each source library; the shaded row is the full library of Table 5 .
Figure 11: Evolved skills versus pretraining references. Each point is one motion: an OMOMO largebox training clip (pink) or one rollout of an evolved skill (colored, 16 per skill; circles mark skill centers). A motion is described by its body posture relative to the pelvis and heading, summarized over time, and embedded with UMAP.
We present MotionDisco, a framework that discovers contact-rich, long-horizon humanoid loco-manipulation motions from scratch, without relying on teleoperation or motion retargeting from human demonstrations. This is challenging because the space of possible contact interactions grows combinatorially with the task horizon and the number of objects in the scene. MotionDisco enables rapid discovery of novel motions by coupling a large language model (LLM) guided evolutionary search over sequences of interactions with an efficient sequential kinodynamic trajectory optimizer and pruning strategy, enabling the rapid discovery of novel skills. Through extensive ablation studies, we show that our LLM-guided search discovers successful whole-body trajectories across several challenging long-horizon tasks. Finally, by training reinforcement learning tracking policies on the discovered trajectories, we transfer the motions to a real humanoid robot. This is the first work to discover and deploy long-horizon humanoid loco-manipulation skills entirely through automated evolutionary search. Supplementary videos of the experiments are available at: https://youtu.be/DHiVz34QYlw.
Ilyass Taouil, Michal Ciebelski, Shafeef Omar +4
Technical University of Munich, Germany · New York University, USA · Carnegie Mellon University, USA
Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their ability to handle complex task sequences. To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagined rollouts of a learned dynamics model. LUCID first trains a structured latent-conditioned low-level policy via adversarial imitation and then freezes it while jointly learning a high-level policy and macro-dynamics world model. The world model predicts the temporally extended state transitions induced by latent decisions, enabling high-level policy optimization through imagined rollouts. We evaluate our framework across various simulated multi-object rearrangement scenarios. Experimental results show that LUCID improves the full-task success and partial-completion rates compared to prior baseline methods, demonstrating its effectiveness in complex sequential loco-manipulation tasks.
Cheng Guo, Mingzhe Ni, Angelo Cangelosi +1
Department of Computer Science, University of Manchester, Manchester, UK · Human-Robot Interfaces and Interaction Laboratory, Italian Institute of Technology, Genoa, Italy
Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion data and a tracking policy improve each other. First, we consolidate motion-captured human-object interaction datasets and retarget them into humanoid robot references while preserving whole-body coordination and dexterous hand-object relationships. This produces a large and diverse humanoid robot reference collection for dexterous whole-body loco-manipulation. Second, we train a physics-based generalist tracker that executes these references in simulation on a humanoid with dexterous hands, covering a scale and diversity beyond prior humanoid tracking systems for loco-manipulation. Third, we close a data flywheel: each round makes small, task-preserving changes to where an interaction takes place and how the body performs it, fine-tunes the tracker on them, and keeps only the variants whose simulated execution completes the task, which seed the next round. With more iterations, these small edits compound into broader coverage around the sparse original demonstrations while preserving task semantics and motion quality. Experiments show contact-preserving retargeting across robot configurations, broad tracking with a single generalist policy, executable motions that keep growing over augmentation rounds, and transfer to real robots. InterMimicGen provides a unified path from heterogeneous human demonstrations to a continually expanding motion resource for humanoid robot learning.