A robot description does more than specify a physical mechanism: it also encodes arbitrary conventions, such as joint-axis direction, joint-angle zero, and the order and names of links and joints. Morphology-aware policies consume interfaces built from these descriptions, yet cross-embodiment evaluation typically changes the robot while keeping those conventions fixed. This leaves a simple question unanswered: does behavior survive when the robot stays fixed but its description changes? GaugeBench isolates this case by rewriting a fixed mechanism under physically equivalent conventions, verifying that its physics and policy interface are preserved, and then evaluating the same policy weights. The result is stark: three MetaMorph policies score 4030.6 on 80 familiar robots, but only 51.6 when those same robots are equivalently re-described, while 98 genuinely held-out robots score 1489.6. A new description can therefore be more damaging than a new robot. Tracing the failure reveals that axis reversal alone reproduces the collapse, joint-angle zero changes are nearly harmless, and reordering lies between them; moreover, changing joint-state and torque coordinates alone is sufficient to cause the failure, while changing description-derived features alone is not. The same phenomenon appears in ModuMorph and an unrelated PyBullet framework. Yet it is not irreversible: exact two-description transport restores the original controller, and training across equivalent axis conventions raises retained return under axis reversal from 3.6% to 80.6%. Together, these results separate mechanism robustness from representation robustness and show that cross-embodiment evaluation should test both.
Figures & tables
Fig. 1: One mechanism, many descriptions, and the evaluation axis that exposes them. A description records a mechanism together with arbitrary conventions: joint-axis directions σj , joint-angle zeros q0,j , link-and-joint order p , and part names. GaugeBench rewrites a robot under a different choice of them, certifies physical and interface equivalence, and only then scores the same weights. MetaMorph keeps roughly a third of its return on a physically held-out robot but almost nothing on an equivalent rewrite of a robot it trained on.
1409
1410
1411
Mean
Reference, original descriptions
100 training robots
4048.8
4168.0
3992.6
4069.8
98 held-out robots
1291.6
1667.7
1509.4
1489.6
80 evaluated robots
4007.9
4145.2
3938.8
4030.6
Equivalent rewrite, same 80 robots
Mean return
51.4
54.3
49.1
51.6
TABLE I: MetaMorph under equivalent rewrites. 3 policies named by training seed, 80 robots, 10 rewrites, 32 episodes per case.
Fig. 2: An equivalent re-description of a familiar robot is more damaging than a physically held-out robot, in all three baseline policy families. Each panel holds one family’s trained policies fixed across the four conditions numbered beneath it: familiar robots under their original descriptions, physically held-out robots, familiar robots under equivalent rewrites, and those rewrites after repair. Only the second changes the mechanism. Bars are performance retained above a do-nothing score, zero for panel (c), averaged over policies, with the absolute score beneath, and points are the individual policies. Transport restores 100.0% aggregate retained performance, matching the original episode for episode except in five of 25,600 paired episodes in one ModuMorph policy. Panel (c) rewrites axis directions alone, because reordering is not certified physically equivalent in that simulator, and its three policies are released artifacts differing in capacity rather than training seeds.
Joint-axis direction
Joint-angle zero
Link-and- joint order
Combined
Rewritten return
146.6
3939.2
1445.9
78.0
Retained (%)
3.6
97.7
35.8
1.9
5th percentile
−17.8
3773.3
612.7
−19.3
Worst case
−26.3
3748.1
528.7
−25.2
Same-robot loss
3884.1
91.5
2584.7
3952.6
95% interval
[3726,4038]
[62,121]
[2434,2733]
[3793,4110]
TABLE II: Effect of individual description conventions.
Fig. 3: The collapse is uniform across robots and policies, and one convention alone reproduces it. (a) Per-robot return on the 80 evaluated robots, sorted by original performance, with bands spanning the three policies. Rewritten return stays near zero across the entire range, so the mean is not carried by a few fragile robots. (b) Per-robot raw ratio of rewritten to original return, one column per trained policy. All 240 robot–policy pairs fall below 15 percent, and the axis is not truncated. (c) Do-nothing-adjusted return retained per convention, against the held-out reference. Joint-axis direction alone reaches the floor the combined rewrite reaches, while the joint-angle zero is nearly harmless. Error bars are the between-policy standard deviation, and these effects are not additive.
Standard
Axis- randomized
Change
Original-description return
4030.6
3606.3
−424
Axis-reversed return
146.6
2911.8
+2765
Return retained
3.6%
80.6%
+77.0 pts
Combined rewrite
1.9%
30.9%
+28.9 pts
Held-out return
1489.6
1491.2
+2
TABLE III: Axis-convention randomization.
Fig. 4: Renaming defeats the trivial fix, but not repair itself. (a) With shared part names, a known correspondence and one recovered from description structure both restore the original return exactly, while correcting observation normalization alone does almost nothing. (b) With every name replaced, name matching recovers nothing and correctly declines to guess, while structure matching recovers all of them, and (c) repair there restores the original return exactly. Bars are means over the three policies, points the individual policies. Every repair is two-sided , receiving both descriptions and retraining nothing.