Co-design methods optimize a robot's body and controller jointly and judge the result by one number, the task reward of the fully optimized pair. That number cannot separate morphologies whose competence depends on the controller to very different degrees. We evaluate a morphology by restricting its controller instead, recording the task competence it retains under an explicitly declared, low-complexity controller family, environment, task and search budget. On three EvoGym locomotion tasks, task reward explains only 39%, 33% and 10% of the variance in this quantity, and geometric descriptors do not predict it under run-grouped cross-validation. The measurement is reliable across optimizer restarts (ICC(2,k)=0.956--0.986) but depends on the declared family: phasing the drive by actuator index instead of position ranks the same morphologies at Spearman 0.50--0.63 and reverses reward-matched pairs. As a second search objective the axis improved competence at matched task reward in 3 of 5 paired runs, short of a pre-registered bar of 4. Used after an ordinary reward-only search instead, to choose within its top task-reward band, it selected a different body in all 8 runs offering a choice, at a cost of at most 0.10 reward units, and in 6 of 8 that body also scored higher under a held-out family. Restricted-control competence is therefore a reportable property of a co-designed morphology, interpretable only with the controller family that defines it.
Figures & tables
Fig. 1: Task reward does not separate robots that depend on their controller very differently. Two co-designed Walker robots, A and B, built from the voxel types listed at the bottom. (a) Top: each robot with the neural-network controller it was co-designed with, which senses the body at every step; the two reach the same task reward to within 0.2% ( 10.58 and 10.60 ). Bottom: the same bodies driven by the declared restricted family instead, one rhythmic wave with four parameters and no sensing; A keeps walking and B barely moves ( 5.93 against 2.12 , mean over three restarts; the rollout shown is the best restart). All frames share one window and scale, and the white number on the ground is the distance traveled in voxels. (b) Distance traveled; the trained rollouts end near step 470 , at the end of the course. Under the ID-phased family the pair reverses ( 1.66 against 2.06 ).
Fig. 2: The protocol, and what must be declared for the number to mean anything. Upper path: what co-design reports, namely the task reward of a controller of unrestricted capacity. Lower path: optimize within a declared restricted family under a declared budget and record the competence that remains. The measured quantity is S(B;C,E,T,Q) , and each of the four conditioning arguments, drawn as the chips reported with every score, changes it. Right: phasing the drive by an actuator’s position rather than its index makes relabeling permute positions and actions together, leaving the action field exactly unchanged, as the inset shows on body A of Fig. 1 .
Fig. 3: The robots and the three tasks. (a) The four voxel types, in the simulator’s colors. At every control step the controller sets each actuator’s target length between 0.6× and 1.6× its rest length along the actuator’s own axis; the resulting deformation against the ground is the only way a body moves. (b) One co-designed robot per task under its trained controller, at four time steps, with the distance traveled so far: walk forward on flat ground, carry the gray box without dropping it, walk down a staircase. Each robot is its task’s highest-reward reward-only design; on Walker, body B of Fig. 1 .
Fig. 4: The measurement is well defined , one panel per way it could fail. (a) Over 40 random relabelings of a fixed body, the ID-phased family changes the commanded action by up to 0.98 , while the position-phased action field is unchanged exactly . (b) Optimizer restarts agree: ICC (2,k)=0.982 , 0.956 and 0.986 for the reported three-restart mean (filled), against 0.916 , 0.813 and 0.932 for a single restart (open); dashed, the pre-registered floor. (c) Median gap between the score at a given budget and at 2Q , as a percentage of the population’s score range; at Q it is 2.6 , 4.9 and 2.4% , inside the pre-registered 5% . (d) Within the middle decile of task reward, task reward spreads by only 0.34 , 0.62 and 0.45 , competence by 2.76 , 1.52 and 2.48 .
Fig. 5: Two defensible restricted families rank the same morphologies differently. (a–c) Percentile rank under the position-phased family against rank under the ID-phased family, for every measured design; Spearman agreement is 0.50 , 0.53 and 0.63 ( n=324 , 274 , 276 ). Gray segments in (a) join Walker morphologies whose task rewards are within 1% of each other and whose ordering reverses between the families; 649 such pairs exist and the 24 largest are drawn. Circles mark the pair of Fig. 1 . (d) That pair: the position-phased family separates the two bodies almost threefold, the ID-phased family puts them the other way round, and the lines cross.
Fig. 6: Sixteen co-designed Walker bodies, sampled evenly across the range of restricted-control competence and ordered by it; voxel colors as in Fig. 3 . Under each body, the upper bar and bold number give its competence, the lower bar and plain number its task reward, on one shared scale. The ordering is not visible in the structure, which is why the score is measured.
Seed
Thr.
MO
n
Reward
n
Δ
0
8.63
5.33
14
4.55
23
+0.78
1
8.62
—
0
3.99
3
empty
2
8.63
3.97
11
3.40
2
+0.57
3
8.61
—
0
4.19
7
empty
4
8.64
2.78
1
2.07
1
+0.72
TABLE I: Primary endpoint of the design experiment: the best reevaluated C2 competence among designs within 5% of the reward-only arm’s best task reward (Thr.: that threshold; n : designs in the band; MO: multi-objective arm). Two runs put nothing in the band.
Run
n
Δ C2
Δ C3
Δ C3 ′
Cost
D, s0 †
23
+2.62
+2.23
+1.52
−0.015
D, s1 †
3
+0.45
+0.65
−0.11
−0.002
D, s2 †
2
+1.19
+1.01
+1.43
−0.048
D, s3 †
7
+0.70
+1.38
+1.72
−0.008
D, s1
3
+0.85
+1.36
+1.20
−0.095
W, s2
12
+2.48
+1.72
+0.67
−0.056
TABLE II: Screening inside a reward-only run’s top 5% task-reward band. Δ is the competence of the most competent band design minus that of the highest-reward one, under the family used to choose (C2), the held-out family (C3), and C3 restricted so that it cannot reduce to C2 (C3 ′ ). Cost is the task reward given up. D: DownStepper, W: Walker. † Reward-only arm of the design experiment.
A robot description does more than specify a physical mechanism: it also encodes arbitrary conventions, such as joint-axis direction, joint-angle zero, and the order and names of links and joints. Morphology-aware policies consume interfaces built from these descriptions, yet cross-embodiment evaluation typically changes the robot while keeping those conventions fixed. This leaves a simple question unanswered: does behavior survive when the robot stays fixed but its description changes? GaugeBench isolates this case by rewriting a fixed mechanism under physically equivalent conventions, verifying that its physics and policy interface are preserved, and then evaluating the same policy weights. The result is stark: three MetaMorph policies score 4030.6 on 80 familiar robots, but only 51.6 when those same robots are equivalently re-described, while 98 genuinely held-out robots score 1489.6. A new description can therefore be more damaging than a new robot. Tracing the failure reveals that axis reversal alone reproduces the collapse, joint-angle zero changes are nearly harmless, and reordering lies between them; moreover, changing joint-state and torque coordinates alone is sufficient to cause the failure, while changing description-derived features alone is not. The same phenomenon appears in ModuMorph and an unrelated PyBullet framework. Yet it is not irreversible: exact two-description transport restores the original controller, and training across equivalent axis conventions raises retained return under axis reversal from 3.6% to 80.6%. Together, these results separate mechanism robustness from representation robustness and show that cross-embodiment evaluation should test both.
Rahath Malladi, Arshia Sangwan, Rajesh K. Gupta +1
University of California San Diego · New York University
Co-design is a high-dimensional search problem in the robot morphology and control design space. Efficient search requires exploiting the structure shaped by their interaction. To understand this structure, we analyze the landscapes of soft locomotion and manipulation tasks. We identify three patterns consistent across regions of their co-design spaces: 1) Within a region, quality varies along a low-dimensional manifold, with minimal variation orthogonal to it, reducing the effective search space dimensionality. 2) In higher-quality regions, the variance in quality is spread across more dimensions, necessitating search to expand dimensionality as quality improves. 3) In higher-quality regions, quality varies along joint morphology-control dimensions, requiring search along them. Using these insights, we devise an efficient co-design algorithm that yields 36% better co-designs than state-of-the-art baselines. We examine their exploration patterns and show that these baselines required an order of magnitude more function evaluations to find co-designs of comparable quality. Finally, we ablate our algorithm to verify that exploiting the identified structure was the key to efficient co-design.
Apoorv Vaish, Oliver Brock
Robotics and Biology Laboratory, Technische Universität Berlin · Science of Intelligence (SCIoI), Cluster of Excellence, Berlin · Robotics Institute Germany (RIG)
An often overlooked factor of robot manipulation performance is the embodiment of the robot itself. Motivated by this problem, we study motion-conditioned robot co-design, where the goal is to generate complete robot designs that track target end-effector trajectories (from human demonstrations) while optimizing user-defined rewards. We introduce Transformer Transformer, a diffusion transformer trained on RoboTokens, a unified tokenization of robot embodiments, states, and actions. The same architecture can be used across embodiment spaces (e.g., wheeled bimanual, quadrupeds, humanoids) and use cases (embodiment generation, cross embodiment controller). Rather than overfitting to one reward function, Transformer Transformer is a dynamics model, whose reward-agnostic state and action predictions can be converted into reward-specific value predictions. These value predictions are used to steer embodiment diffusion towards high value robot designs, through a procedure we call Dynamics Self-Guidance. Experiments across multiple design spaces show zero-shot optimization of unseen rewards and trajectories, improving performance and runtime over the evolutionary baseline. Finally, we fabricated an optimized ALOHA design, which reduced tracking error by over 70% compared to the original design.