Draft: A Parametric Tool for Robot Design Exploration
Authors: David Nguyen, Marcelo Coelho, Sangbae Kim
Organizations: Department of Mechanical Engineering, Massachusetts Institute of Technology · Department of Architecture, Massachusetts Institute of Technology
Robot performance is often limited by the cost of iterating on morphology and control together, since every computer-aided design (CAD) change has to be carried into a simulation-ready model before control work begins. Co-design methods attempt to close this gap, but each uses a model generator written for a single platform or lack the use of real-world data to suggest that designs are plausible. We present Draft, a parametric generation tool whose generalized engine compiles any parametric tree of serial chains into a simulation-ready MJCF model, without CAD. It allows engineers to explore design tradeoffs through easily adjustable models and evaluate how changes influence controller performance. Draft grounds the free parameters of each design using trends fitted to a survey of 114 actuators and 49 published robot descriptions, so that a generated robot is anchored to real-world hardware. We validate those trends wholistically by building twins of four off-the-shelf robots, whose masses agree to 1.10× geometric mean fold error. Finally, we demonstrate how Draft exposes design tradeoffs by evaluating three quadrupeds through a two-stage reinforcement learning curriculum.
Figures & tables
Fig. 1: Draft compiles medium-fidelity robot models from a small set of specifications into a simulation-ready MJCF model. One parameter edit propagates through visual, collision, inertial, and actuation properties.
Fig. 2: All generated Draft models are constructed from four main link primitives. The basic link includes a link and motor at its distal end, the sphere link terminates articulated chains with a contact point, the foot link does the same but with a support surface, and the root link holds custom geometry for limbs to extend from.
Fig. 3: An actuator’s parameters are derived from data-informed relations. The inputs are peak torque τ , no-load speed ωNL , and aspect ratio a . Mass m is estimated from τ , and gear ratio N from ωNL . Volume V is estimated from τ and N ; together with a , it determines radius r and length ℓ . Finally, the rotor inertia J is derived from r .
Fig. 4: The four relations that dictate gear ratio, mass, volume, and rotor inertia plotted against the raw data. Panel (a) shows the power-law fit for gear ratio against no-load actuator speed, while (b) shows the fit for actuator mass against peak output torque. Panels (c) and (d) show prediction performance for volume and rotor inertia, respectively. Points are grouped by actuator class (QDD, MidGear, and High GR) and show agreement with the fitted predictions (black lines). The blue bands are ±25% in (c) and (d).
Parameter
Fitted value
n
R2
LOO
Reduction N
327ωNL−1.066
103
0.89
1.33×
Mass m
0.0653τ0.666
103
0.90
1.28×
Volume V
τ0.700N−0.167/23948
103
0.87
1.30×
Radius r
(V/2πa)1/3
103
—
1.12×
Length ℓ
2ar
103
—
1.14×
Inertia J
27.2r4
88
0.84
1.56×
TABLE I: Fitted trends in pipeline order, with LOO fold error.
Fig. 5: The generation pipeline run end to end against four off-the-shelf robots. The twin takes the vendor’s geometry and declared torque and speed, and derives all other parameters. On the left are renders of the vendor and twin models with the twin’s mass ratio as a percentage above. On the right is the dynamics over 300 configurations per pair with inertia above and gravity below. The second spread for humanoids only holds the misaligned axes still.
Humanoid
Quadruped
Class
n
c
e
Fold
n
c
e
Fold
Thigh
54
13.8
1.36
1.38
60
28.3
2.69
1.43
Shank
50
4.11
1.68
2.09
60
9.25
2.38
1.21
Hip link
56
2.58
1.05
3.18
72
0.60
0.51
2.88
Head
16
2.50
0.63
1.74
—
—
—
—
Upper arm
30
9.40
1.83
1.88
—
—
—
—
TABLE II: Structural mass by segment class, fit as m=cLe with L in meters, and its in-sample fold error.
Inertia
Gravity
Swept
Axes Held
Swept
Axes Held
Robot
Med.
Var.
Med.
Var.
Med.
Var.
Med.
Var.
(%)
(% 2 )
(%)
(% 2 )
(%)
(% 2 )
(%)
(% 2 )
G1
33.9
106
28.3
72.9
25.1
173
19.5
73.6
H2
32.6
103
21.9
47.0
34.1
249
14.5
48.2
Go2
12.8
2.4
—
—
21.4
8.4
—
—
TABLE III: Twin error in M(q) and g(q) over 300 poses, median and variance, with misaligned axes held.
Fig. 6: Three quadrupeds from one specification, differing in scale, leg length, and actuators, rendered in one scene at true relative size.
Fig. 7: The same curriculum is used for each of the three robots. A base policy is trained over a small ranges of commanded speeds vcmd , terrain heights zt , and pushes Δv . Next, three task policies train the base policy further, each ramping one of axis. Every policy is trained from ten seeds.
Fig. 8: Four capability axes for the designs of Fig. 6 , in absolute units. Lines are medians over ten training seeds and bands span the seed minimum to maximum; dashed lines are the same policies on the builds carrying the inertia error of Sec. IV-E . (a) Commanded against achieved speed, hollow past each design’s peak. (b) Traversal, ascent and descent pooled. (c) Recovery from a velocity kick. (d) Drivetrain efficiency up to each design’s peak.
Term
Weight
Shaping
Velocity tracking
1.0
σ=0.5s m/s
Yaw-rate tracking
1.0
σ=0.5/s rad/s
Upright
1.0
σ=0.45/s
Air time
0.4
tmax=0.35s s
Base height
0.4
h∗=0.43s , σ=0.11s m
Foot clearance
−2.0s−3/2
hf∗=0.10s m
TABLE IV: Stage-one reward terms; s enters only where shown.
Speed
Step
Kick
Design
kg
L (m)
m/s
v^
m
L ’s
m/s
v^
η
Cheetah
33
0.48
3.63
1.67
0.12
0.26
7.9
3.65
0.20
Bear
110
0.74
4.33
1.61
0.12
0.17
8.6
3.20
0.47
Giraffe
91
1.30
2.17
0.61
0.20
0.15
5.3
1.48
0.41
TABLE V: Capability, absolute then per leg length L ; v^=v/gL . Medians over ten seeds; bold is best.
Sprint
Kick
Terrain
Design
τ %
ω %
τ+ω
τ %
ω %
τ+ω
τ %
ω %
τ+ω
Cheetah
K
95
25
1.05
H
69
17
0.60
K
43
13
0.44
Bear
K
77
85
1.04
H
66
59
0.93
K
43
50
0.55
Giraffe
H
96
105
1.08
H
94
99
1.06
H
76
71
1.00
TABLE VI: Utilization at the limiting joint, knee (K) or hip pitch (H); bold is actuator limited.
We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in physics-based evaluation. A design agent and control agent engage in a thesis-antithesis-synthesis loop, while criterion-specific LLM judges provide multi-objective feedback to steer exploration. Across five MuJoCo locomotion benchmarks, D2C achieves the highest default-normalized score among the evaluated LLM-based and black-box baselines, with gains up to 3.2x on Ant and nearly 9x on Swimmer. Iterative debate yields 18-35% gains over compute-matched zero-shot generation, and D2C-generated rewards transfer to default morphologies in 4/5 tasks. These results suggest that structured, simulator-grounded multi-agent interaction is a useful mechanism for joint morphology-reward optimization under a fixed-topology, per-candidate-RL protocol. Project page: debate2create.github.io.
An often overlooked factor of robot manipulation performance is the embodiment of the robot itself. Motivated by this problem, we study motion-conditioned robot co-design, where the goal is to generate complete robot designs that track target end-effector trajectories (from human demonstrations) while optimizing user-defined rewards. We introduce Transformer Transformer, a diffusion transformer trained on RoboTokens, a unified tokenization of robot embodiments, states, and actions. The same architecture can be used across embodiment spaces (e.g., wheeled bimanual, quadrupeds, humanoids) and use cases (embodiment generation, cross embodiment controller). Rather than overfitting to one reward function, Transformer Transformer is a dynamics model, whose reward-agnostic state and action predictions can be converted into reward-specific value predictions. These value predictions are used to steer embodiment diffusion towards high value robot designs, through a procedure we call Dynamics Self-Guidance. Experiments across multiple design spaces show zero-shot optimization of unseen rewards and trajectories, improving performance and runtime over the evolutionary baseline. Finally, we fabricated an optimized ALOHA design, which reduced tracking error by over 70% compared to the original design.
We propose to turn generalist multi-embodiment value functions into reusable models for robot design. Instead of running a new reinforcement learning co-design loop for each robot, we first train an embodiment-aware policy and value function across many robot designs. After training, the frozen value function is used as a differentiable surrogate to optimize candidate embodiments through value gradients. We evaluate our approach across different robot design settings, from perturbed single robots to held-out robots across morphology classes, with single models trained on up to 50 robots and design spaces of over 1100 continuous embodiment parameters. Beyond optimizing complete embodiments, we show that value gradients can identify performance-limiting design and control parameters, enabling both the optimization and the analysis of new robot designs.
Nico Bohlinger, Jan Peters
Technical University of Darmstadt, Germany · Robotics Institute Germany (RIG); German Research Center for AI (DFKI); hessian.AI