Robot co-design couples morphology search with policy learning, yet training every new design from scratch discards acquired control experience. We present LACE-CRAFT, which compares continued learning on the current robot with policy adaptation to new morphology-reward pairs. LACE resumes the incumbent's full learning state and initializes compatible challengers with its actor parameters and observation statistics. A fixed task metric selects among both branches and the frozen incumbent. CRAFT coordinates Feedback, Morphology, Reward, and Integration roles through shared experimental records and behavioral replays to generate and cross-review paired proposals. A generative extension converts generated meshes into editable articulated models with configured joints, actuator interfaces, and consistently updated simulation assets. Across five locomotion benchmarks, mean scores over three evaluation seeds are 6.4-91.9% higher than D2C. Both methods train 30 new morphology-reward pairs over five rounds; LACE additionally uses four continuation training units. Five-task ablations examine policy inheritance and replay-derived feedback. A fabricated prototype demonstrates indoor walking and illustrates the geometry-to-hardware workflow.
Figures & tables
Component
Continuation
Challenger
Morphology and reward
Unchanged
New proposal
Actor parameters
Restore
Copy
Observation statistics
Restore
Copy
Value / critic networks
Restore
Initialize anew
Optimizer states
Restore
Initialize anew
SAC targets and temperature
Restore
Initialize anew
TABLE I: Training-state handling before each branch learns. Copied Actor parameters and observation statistics continue to update.
Record
Shared contents
Experiment history
Candidate identifier, selected score, and linked morphology, reward, and checkpoint files.
Current evidence
Saved replay video, trajectory and task measurements, and the selected-checkpoint reference.
Role
Output
Feedback
Failure hypotheses.
Morphology
Body edits and reward cross-review.
Reward
Reward code and body cross-review.
TABLE II: Shared CRAFT evidence and role-specific outputs.
Fig. 2: An input image and user requirement define editable morphology m through Rodin reconstruction and BANG decomposition. The four views show complementary components, not sequential steps. The base link incorporates an SO-101 mount, while other arm components reuse the SO-101 design. CRAFT supplies morphology edits and training rewards. LACE receives the compiled model and selects trained candidates using a task metric distinct from training reward. Replays and scores return to CRAFT. Simulation evaluates walking with fixed arm postures, not manipulation. Physical walking excludes the arm. The mounted-arm photograph documents assembly only.
Fig. 3: Three-seed evaluation means with ±1 sample-standard-deviation error bars. Means and deviations are divided by the task’s fixed D2C mean; denominator uncertainty is not propagated.
Task
Default
D2C
LACE–CRAFT
Ant
27,362.80 ±621.94
33,044.51 ±477.05
63,396.82 ±656.01
HalfCheetah
18,343.94 ±201.20
20,072.15 ±18.82
27,455.51 ±22.39
Hopper
4,377.53 ±9.80
6,195.22 ±2.29
7,125.95 ±11.94
Swimmer
897.83 ±8.19
6,637.06 ±10.67
9,733.43 ±68.55
Walker2D
10,114.19 ±10.37
17,362.33 ±221.69
18,466.68 ±29.89
TABLE III: Final-policy evaluation scores: mean with ± sample standard deviation below, over three evaluation seeds (128 episodes each).
Fig. 4: Three-seed evaluation means with ±1 sample-standard-deviation error bars, divided by the task’s fixed full-method mean. Denominator uncertainty is not propagated.
Task
Full
No Actor
No replay
Ant
63,396.82 ±656.01
54,806.01 ±432.08
50,609.18 ±838.12
HalfCheetah
27,455.51 ±22.39
25,147.09 ±116.61
25,272.34 ±25.86
Hopper
7,125.95 ±11.94
6,600.18 ±9.61
6,693.74 ±2.87
Swimmer
9,733.43 ±68.55
4,278.96 ±23.92
4,832.12 ±2.60
Walker2D
18,466.68 ±29.89
10,919.33 ±133.34
13,171.87 ±232.93
TABLE IV: Final-policy ablations: mean with ± sample standard deviation below, over three evaluation seeds (128 episodes each).
Fig. 5: Hardware-adapted printable parts and assembly with the SO-101 arm.
Fig. 6: Indoor walking with the arm removed. Frames are from one video; timestamps are in seconds.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Continuation
Challenger
Morphology and reward
Unchanged
Proposed pair
Actor parameters
Restore
Copy
Observation statistics
Restore
Copy
Value / critic parameters
Restore
Initialize anew
Optimizer states
Restore
Initialize anew
SAC targets and entropy state
Restore
Initialize anew
Appendix
TABLE S1: Checkpoint handling for the current LACE protocol.
Record
Stored content and use
Selected robot
Candidate and checkpoint identifiers, morphology and reward bindings, score, and full-state checkpoint.
Selection history Ak
Compared candidates, scores, and round-specific selection flags; the bounded view enters the next discussion.
Evidence Ek
Selected-checkpoint rollout summary and references to replay and visual manifests.
Role archive Hk
Requests, responses, drafts, revisions, integration rationales, validation errors, and API records; retained separately from future role inputs.
Appendix
TABLE S2: Persistent records and their use in CRAFT.
Setting
Ant
HC
Swimmer
Learning rate
3×10−4
3×10−4
3×10−4
Discount
0.97
0.95
0.97
Entropy cost
0.01
0.001
0.001
Reward scaling
25
3
5
Training environments
256
128
128
Batch size
512
512
128
Appendix
TABLE S3: PPO configuration. HC denotes HalfCheetah.
Setting
Hopper
Walker2D
Learning rate
6×10−4
6×10−4
Discount
0.997
0.997
Reward scaling
30
5
Training environments
12
12
Batch size
48
48
Gradient updates per step
3
3
Appendix
TABLE S4: SAC configuration.
Task
RL
Requested
Actual
Horizon
Ant
PPO
8,000,000
8,007,680
450
HalfCheetah
PPO
5,000,000
5,038,080
400
Hopper
SAC
400,000
400,008
800
Swimmer
PPO
350,000
350,208
600
Walker2D
SAC
500,000
500,004
1,000
Appendix
TABLE S5: Task-specific learning configurations. Actual transition counts include batch alignment.
TABLE S7: Initialization of the reported five-task searches.
Fig. S1: Ant candidate scores in rounds 2–4 and the five-round selection path. Stars mark winners. Inc.: incumbent; Cont.: continuation; C1–C6: challengers ordered by identifier within each round.
Search
RL seed
Selected score
1
0
57,560.79
2
0
55,259.83
3
0
64,252.91
Appendix
TABLE S8: Additional Full searches on Ant; 30 generation units per search. Scores are historical selection scores.
Fig. S2: Exploratory search-level variation on Ant versus cumulative generation units. Full shows the mean and sample standard deviation of three searches at RL seed 0; other methods have one search each.
Round
D2C
LACE–CRAFT
Gain
1
11,023.11
12,435.48
+12.8%
2
14,422.04
16,488.11
+14.3%
3
14,902.87
16,742.91
+12.3%
4
13,902.85
15,088.00
+8.5%
5
13,186.34
16,394.60
+24.3%
Appendix
TABLE S9: Walker2D: best new design in each round.
Task
Variant
Score
Units
Cont.
Ant
Full
64,252.91
34
4
No Actor
56,901.25
34
4
No replay
51,878.46
34
4
HalfCheetah
Full
27,458.55
34
4
No Actor
25,263.65
34
4
No replay
25,277.06
34
4
Appendix
TABLE S10: Historical search-selection scores and completed training budgets. These are not the post-selection three-seed means in the main text. Units include continuations.
Task
Variant
Round 1
Round 2
Round 3
Round 4
Round 5
Ant
No Actor
34,631.61
50,914.84
50,914.84
56,901.25
56,901.25
No replay
29,049.02
42,973.31
48,472.22
51,878.46
51,878.46
HalfCheetah
No Actor
25,263.65
25,263.65
25,263.65
25,263.65
25,263.65
No replay
20,733.75
23,762.41
24,715.28
25,277.06
25,277.06
Hopper
No Actor
5,212.28
6,595.88
6,595.88
6,595.88
6,595.88
No replay
6,573.45
6,573.45
6,573.45
6,573.45
6,693.28
Appendix
TABLE S11: Selected scores after each ablation round. Repeated values indicate retention of the best available robot.
Stage
Output
Shape generation
Static mesh conditioned on the morphology reference.
Decomposition and alignment
Separate link meshes gi with local frames and aligned connections.
Joint allocation
Parent–child graph, joint origins, types, axes, and position limits.
Actuator fitting
Motor positions and orientations, cavity geometry, and mating surfaces in Σ .
Compilation
URDF, collision geometry, inertial properties, joint and actuator limits, and action mapping in Φ .
Appendix
TABLE S12: Artifacts produced by generative-morphology processing.
Fig. S3: Ant single-search progress versus cumulative generation units. Each curve shows stored selection scores from one search, not the post-selection evaluation means.
We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in physics-based evaluation. A design agent and control agent engage in a thesis-antithesis-synthesis loop, while criterion-specific LLM judges provide multi-objective feedback to steer exploration. Across five MuJoCo locomotion benchmarks, D2C achieves the highest default-normalized score among the evaluated LLM-based and black-box baselines, with gains up to 3.2x on Ant and nearly 9x on Swimmer. Iterative debate yields 18-35% gains over compute-matched zero-shot generation, and D2C-generated rewards transfer to default morphologies in 4/5 tasks. These results suggest that structured, simulator-grounded multi-agent interaction is a useful mechanism for joint morphology-reward optimization under a fixed-topology, per-candidate-RL protocol. Project page: debate2create.github.io.
In this paper, we introduce a model of evolution and learning in robots that co-optimizes a distribution of latent design vectors (genotypes) and a mixture of control experts (neural modules), which are gated by the latent coordinates of each decoded design (phenotype). This provides a scalable alternative to co-design algorithms that either train an individual policy for every robot, which is inefficient, or a monolithic universal controller for all robots, which results in overly conservative structures and behaviors. Our approach lies somewhere between these two extremes, preserving ancestral knowledge in a unified yet modular framework in which different body plans activate and deactivate different combinations of learned sensorimotor circuits for goal-directed behavior. This allows one part of the controller to be overhauled to better suit new species of designs as they emerge without disrupting the hard-earned knowledge contained within other expert modules. It also allows pretrained expert policies to be directly plugged into the mixture, which can steer evolution into otherwise unexplored areas of latent space containing desired morphological traits. We refer to this process as "evo by demo" and explore how it may be used to guide freeform evolution toward canonical structures defined by the pretrained model. Videos and code can be found at: https://eco-moe.github.io.
Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts inside COACHWORLD, our shared action-conditioned world model, and uses a progress judge to record the first subtask that fails to complete. Aggregated records select which subtask demonstrations to acquire and which expert adapters to update. Across two simulation suites and two real-robot platforms, imagined and deployed success correlate over 22 task-policy pairs (rho = 0.840). Controlled comparisons show that our coaching method outperforms matched baselines under matched data budgets and update schedules. With only 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The coached experts also transfer to four held-out compositions, achieving an average success of 35.0%, compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. Together, these results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://robocoach-ai.github.io/
Jiajun Liu, Yifan Chen, Yichao Liu +7
Renmin University of China · Tsinghua University · Shanghai Qizhi Institute +2