We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.
Figures & tables
Figure 1: InterEvolve adapts a frozen whole-body controller with novel use. The controller trained on human-object interaction (HOI), such as lifting and pushing a box, is asked to rotate a box around in place, a behavior absent from its training data (Sec. D.7 and Figure 11 ). A human-designed reward tries to hack this behavior but fails. The reward program that InterEvolve evolves compensate with novel body used to succeed. Bottom: InterEvolve can achieve versatile evolved skills in simulation and on a real robot. More demos are on the project page .
Figure 2: InterEvolve overview. (a) Given a task and scene, an LLM agent writes a staged reward program, drawing on verified programs in the skill library. (b) The forward-backward model scores the active stage’s reward over a fixed state bank and projects it into a latent prompt, which a frozen body prior and a pretrained object residual execute without retraining. (c) A fixed verifier scores parallel rollouts, and criterion-level feedback drives the next structural revision (outer loop), while CMA-ES tunes the program constants (inner loop).
Table 3
Figure 3: Dexterous hands and long-horizon composition. (a) The G1 with Inspire hands, tracking a reference with the tracking reward calibrated by InterEvolve. Left : whole-body view at the moment of lift; Right : close-ups of the manipulation phases. (b) Arranging six boxes by the evolution. Left : top-down trace of the robot (gray) and each box from its start (outlined square) to its target cell (dashed circle); numbers give the push order. Right : the robot pushing each box.
Search components
Results
CMA-ES tuning
Targeted edits
Multi- scenario
Scene context
Multi- stage
SR (%) ↑
Earned (%) ↑
GPU-h ↓
Tokens (M) ↓
✓
✓
✓
✓
✓
86.5
95.6
2.1
0.23
✗
✓
✓
✓
✓
51.6
89.9
1.7
0.27
✓
✗
✓
✓
✓
68.9
92.3
2.7
0.20
✓
✓
✗
✓
✓
68.4
89.9
1.0
0.17
✓
✓
✓
✗
✓
78.7
93.2
2.0
0.21
Table 3: Ablation of the search design. Stages and numerical calibration are most important designs. Each row disables one component (✗) of the full system; definitions are in Sec. D.5 .
Figure 4: Autonomous execution of real-world G1 with onboard camera and task intent of kicking a box (top) and pushing a box (bottom). Each row runs a program evolved in simulation. The first four columns show the run from a third-person view, left to right in time. The last two columns are the robot’s onboard color and depth images as it approaches the box, overlaid with the FoundationPose estimate (yellow) and the filtered pose that the controller receives (green).
Table 7
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Interaction encoding
Dim.
Eh↓
Eo↓
SR ↑
Δ SR
Object state + link geometry
105
27.20
30.80
60
n/a
No object state
90
28.57
49.20
20
−40
No link geometry
15
25.05
60.83
12
−48
Appendix
Table 6: Interaction observations. Ablations remove the object state or the link geometry. Δ SR is relative to the full 105-dimensional encoding, which is the main model of Table 2 (without replanning).
Model
Trainable parameters
Eh↓
Eo↓
SR ↑
From scratch
843 M
60.96
78.21
2
Residual on the body prior
MLP 256×2
14.1 M
36.19
50.81
28
MLP 512×2
30.5 M
31.51
42.58
30
MLP 1024×2
79.5 M
29.22
37.42
46
Full-width residual
852.0 M
27.20
30.80
60
Appendix
Table 7: Residual capacity. MLP width × depth of the residual branch; the full-width residual mirrors the prior’s module architectures and is the main model of Table 2 (without replanning). Parameter counts include all trainable modules. Errors and SR follow Table 2 .
Anchor
Addition
kp
SR (%) ↑
Eh↓
Eo↓
None
–
–
60
27.20
30.80
Anchor
Root
–
2
68
19.56
25.27
Object
–
3
72
19.80
24.92
Root and object
–
2, 3
68
18.24
24.10
Correction gain kp
Appendix
Table 8: Drift-correction ablation. Tracking protocol of Table 2 ; Eh and Eo in cm. The anchor is the quantity steered toward the reference, kp is the correction gain, and the shaded row is the configuration used in Table 2 .
Figure 5: World-frame drift over time. Horizontal distance between the simulated and reference root (left) and object (right) for native tracking and the object-anchored drift correction, on the 50 largebox clips of Table 2 . Lines show the median and bands the 25–75% range over all 50 motions; each motion is counted to the end of its clip, and after its clip ends its final value is held. Without correction the median root drift grows to 0.29 m at 4 s and 0.32 m by the end, against 0.16 m and 0.18 m with the correction, which levels off after about 4 s. The two curves coincide for the first second because the correction has a 0.10 m deadband.
β
State count
SR (%) ↑
N90
Projection latency (ms) ↓
Peak memory (MiB) ↓
Reward tilt
0 (uniform)
50k
0.0
45,000
–
–
1
50k
15.6
28,674
–
–
3
50k
35.9
10,655
–
–
10
50k
59.4
2,077
0.92
63
30
50k
28.1
3
–
–
Appendix
Table 9: Reward-inference bank diagnostics on the feet-only task. SR is the task success rate of the fixed program; N90 is the number of bank states that carry 90% of the weight at a frame, averaged over the frames of a rollout and reported as the median over the 64 test scenarios. Latency is per reward projection; peak memory covers resident bank features and projection workspace.
Figure 6: Reward tilt, bank size, and cost on the feet-only task. (a) SR against the reward tilt β with 50k states. (b) SR against bank size at β=10 . (c) Peak memory (solid) and projection latency (dashed) against bank size.
Component
Symbol
Role
Example edit
Reward code
rj
Desired body-object interaction
Upward → horizontal object motion
Constants
θ
Scale and target of an objective
Target height, contact offset
Stages
j
Intermediate objectives
Split acquisition from transport
Completion
gj
Completing or switching a stage
Require support before transport
Stage memory
–
Context captured at runtime
Displacement since stage entry
Appendix
Table 10: Editable components of a reward program. Each edit is judged by its executed outcome.
Figure 7: Reward-program evolution on the kick task. Each row shows the current best program after the initial proposal, rounds 1–3, and the final round (top to bottom), the kick programs evaluated in Table 5 ; frames run left to right. Revised objectives change the contact strategy across rounds.
Figure 8: The eight basic skills. Final evolved program of each task family, frames left to right. Whole-body motion links object acquisition, transport, and placement across carrying, lifting, tipping, kicking, and obstacle-aware pushing.
Figure 9: Composite tasks. The three composite tasks of Table 5 : relocating a box between supports, stacking one box on another, and carrying, placing, then kicking a box. Frames run left to right.
ULTRA
BFM-Zero
InterEvolve
+ Evolving
SR (%) ↑
66
8
60
72
Eh (cm) ↓
15.68
28.36
27.20
19.80
Eo (cm) ↓
21.45
62.33
30.80
24.92
Orientation error ( ∘ ) ↓
37.8
47.6
23.4
26.2
Contact retention (%) ↑
92
2
89
89
Slip, sim / ref. (cm/s)
27.7 / 19.4
20.0 / 26.4
20.0 / 20.8
19.6 / 20.5
Appendix
Table 11: Supplementary tracking metrics on the clips of Table 2 . Every metric is averaged over three evaluation runs of the 50 clips, with rates over clips rounded to integers, and the rows below the first block come from recordings in which every clip runs to its end without early termination. Orientation error : object orientation error, averaged over the whole clip. Contact retention : fraction of reference-contact frames in which the simulated hand is also in contact, with contact meaning a hand vertex within 3 cm of the object surface. Slip : wrist speed in the object frame during contact, in simulation and on the reference. Penetration , object below ground : frames deeper than 2 cm. Falls : pelvis below 0.35 m or tilted beyond 60∘ .
Task
Interaction
Key constraint
Push to mark
Push object to a ground target
Keep it on the floor, stop near target
Carry at mid height
Acquire, raise, and carry
Hold a waist-relative height band
Carry at chest height
Acquire, raise, and carry
Hold a body-relative height
Tip onto a new face
Turn it over about a horizontal axis
Large tilt, left standing on the new face
Push through a gate
Push between two immovable walls
Hold a straight heading through the gap
Push around an obstacle
Push past a pillar in the way
Deviate, then recover the line
Appendix
Table 12: Task families for program search. The verifier scores outcomes and constraints independently of the agent-written reward.
Task family
Human + CMA
Initial agent + CMA
InterEvolve
Push to mark
96.9
100.0
100.0
Carry at mid height
0.0
40.6
100.0
Carry at chest height
0.0
0.0
100.0
Tip onto a new face
0.0
0.0
95.3
Push through a gate
0.0
25.0
73.4
Push around an obstacle
0.0
4.7
71.9
Appendix
Table 13: Success per task family (%). CMA denotes CMA-ES. The macro-average weights families equally.
Figure 10: Push to mark under new goals and starts , with the selected program unchanged. (a) From one start, targets at bearings from −45∘ to +45∘ , 2.5 m away; rings mark the targets. (b) The box starts 0.55, 1.2, or 1.8 m from the robot at bearings of −40∘ , 0∘ , and +40∘ , and is pushed to one target (red ring). Lines trace the box; each rollout shown is one of 16 per setting.
Task family
Large box
Plastic box
Small box
Suitcase
Push to mark
100.0
0.0 (4.7) → 98.4
0.0 (0.0) → 98.4
0.0 (25.0) → 92.2
Carry at mid height
100.0
15.6 → 84.4
35.9 → 96.9
76.6 → 90.6
Carry at chest height
100.0
18.8 → 98.4
29.7 → 95.3
98.4 †
Tip onto a new face
95.3
78.1 → 93.8
28.1 → 89.1
31.2 → 96.9
Push through a gate
73.4
40.6 → 98.4
56.2 → 85.9
68.8 → 90.6
Push around an obstacle
71.9
32.8 †
1.6 → 84.4
96.9 †
Appendix
Table 14: Selected programs on other boxes. Each program of Table 13 is searched on the large box and run on the three other box-like OMOMO objects the controller was trained on, first unchanged (left of the arrow) and then adapted to each box by hand; † marks boxes where adaptation did not help and the unchanged program is kept. SR (%) over 64 test scenarios, as in Table 13 ; the large-box column is a rerun of the same protocol and reproduces that table exactly.
Task
Scene and interaction
Verifier criteria
Relocate between supports
The box starts on a 0.3 m support in front of the robot. The robot lifts it, turns around, carries it 2.9 m to a 0.6 m support behind it, sets it down, and lets go.
Stays upright; holds the box for a large part of the episode; lifts it above the destination height; ends within 0.3 m of the mark on the support; hands clear of the box at the end.
Stack two boxes
Box A is on the floor 0.5 m ahead; box B, of the same size and free to move, is 1.9 m to the left. The robot lifts A, carries it above B, lowers it onto B’s top face, and lets go.
Stays upright; lifts A clear of B’s top; sustained two-handed contact; A stays level; ends within 0.25 m of B’s top (the target follows B if B is pushed); hands clear; A at rest.
Carry, place, and kick
The box is on the floor 0.5 m ahead. The robot pushes it about 1.0 m along the ground with its hands, lets go, and kicks it with its feet a further 1.5 m to a mark 3.0 m ahead.
Stays upright; box stays low (no lift); sustained hand contact during the push; at least one foot contact; box travels at least 1.0 m after the last hand contact; hands clear; ends within 0.3 m of the mark.
Appendix
Table 15: Composite tasks for library reuse. Distances are measured from the box’s starting position.
Stored program, no search
SR (%)
Kicking to mark
6.3
Push to mark
0.0
Carry at mid height
0.0
Carry at chest height
0.0
Push around an obstacle
0.0
Push through a gate
0.0
Appendix
Table 16: Library ablations on carry, place, and kick , extending Table 5 . Every row is scored on the same 64 test scenarios, none of which any search used; SR is the percentage of scenarios solved, where a scenario is solved when one of three rollouts meets every criterion. Left: each stored program run as written, without search. Right: three search rounds from each source library; the shaded row is the full library of Table 5 .
Figure 11: Evolved skills versus pretraining references. Each point is one motion: an OMOMO largebox training clip (pink) or one rollout of an evolved skill (colored, 16 per skill; circles mark skill centers). A motion is described by its body posture relative to the pelvis and heading, summarized over time, and embedded with UMAP.
Department of Computer Science, University of Manchester, Manchester, UK · Human-Robot Interfaces and Interaction Laboratory, Italian Institute of Technology, Genoa, Italy