Modern methods reconstruct or generate simulation-ready articulated objects, predicting not only their geometry but also how their parts are connected and allowed to move. Evaluating the geometry is straightforward, but evaluating the predicted articulation is not, because articulation specifies a motion rather than a shape, and there is no agreed distance between two motions. More specifically, existing protocols score joint type, axis direction, origin, and motion limits separately, although these parameters jointly describe a single physical motion, and the same motion can be written as different parameter values. As a result, a joint can score maximally wrong against an equivalent encoding of itself, and several component errors are ill-conditioned or undefined exactly where predictions become accurate. We propose ArticulateArena, a representation-invariant counterpart of Chamfer distance for articulation that compares the motions one-DOF joints induce rather than the parameters that encode them. It represents each joint by the unordered pair of its Lie-algebra endpoint twists, and we prove that the resulting quotient distance is a metric. It unifies fixed, revolute, prismatic, and helical joints, brings continuous joints into the same score through a compactification, and reads as the RMS motion of the moving part in meters when weighted by its mass distribution. A motion-aware tree edit distance lifts the metric to full kinematic trees, pricing structural errors such as spurious or missing joints in the same motion units as joint errors, and for a fixed inner product it remains a metric on trees up to relabeling. Alongside the metric we release ArticulateArena-20K, a new suite of 19,977 articulated objects with verified kinematics, and we re-evaluate published reconstruction methods on it under the new metric. Project page: https://heyumeng.com/ArticulateArena-web/
Figures & tables
Figure 1: Categories of joints in articulated objects: fixed, prismatic, revolute, continuous, and helical. Each panel shows an object with one part displaced along its joint (dashed box), an inset of the same part at rest, the kinematic tree with that joint highlighted, and an icon of the motion the joint allows. These five one-DOF categories are exactly the joints our metric covers.
Symbol
Meaning
J
a one-DOF joint
t
joint type
a
joint axis (unit)
o
joint origin
l=[l−,l+]
limit interval
q
joint coordinate (rad or m)
Table 1: Notation used throughout the paper.
Figure 2: From a URDF joint to our metric. (1) A joint is specified by type, axis, origin, and limits. (2) The twist ξ=(ω,v) absorbs type, axis, and origin, and the limits give the endpoint pair z±=l±ξ . (3) E compares two endpoint pairs under the better of the direct and the swapped assignment. (4) Radial compactification maps unbounded ranges into the unit ball, so continuous joints become boundary points.
Finite
Unbounded
Trees
Split norm ∥⋅∥α
Eα
Eαϕ
Eαϕ,tree
Kinetic norm ∥⋅∥B
EB
EBϕ
EBϕ,tree
Table 2: The design space. An inner product (rows) combined with a limit treatment or the tree lift (columns). Shaded entries mark the reported variant.
Figure 3: Overview of ArticulateArena-20K . (a) One object from each of the 15 supercategories, rendered with its kinematic tree (top right). (b) The distribution of the library’s joints over the five joint types. (c) Objects and categories per supercategory, stacked by source library.
Case
Per-component score
Ours
1
Axis reversed, limits negated
elimdir=2 , maximal against itself
E=0 , endpoints swap
2
Interval shifted, same width
elim=0 although the configurations differ
E>0
3
Near-parallel axes
eorigrev divides by ∥apred×agt∥→0 , undefined at parallel
no denominator, shrinks continuously
4
Prismatic origin moved along the part
eorigpris>0
E=0 , origin absent from the twist
5
Revolute range [0,ϵ] against fixed
etype=1 for every ϵ>0
E=ϵ∥ξ∥→0
6
Continuous joint
elim undefined, m infinite
boundary point under ϕ , finite ranges approach it
Table 3: Failure cases of per-component scores. Rows 1 and 4 compare two encodings of the same motion, row 2 two different motions that a component score cannot separate, and rows 3 and 5 two motions that converge to each other.
Figure 4: Calibration of Eα (left) and EB (right) against the true material-motion discrepancy d(g1,g2) of the child link, computed on the actual SE(3) displacements rather than on their linearization. Axis, origin, and limit perturbations (colors) of 690 joints from 60 objects whose child links span a wide range of sizes. The dotted line is the least-squares fit in log space.
Figure 5: Motion-aware edit costs ( Eαtree computed exactly, amber, median and interquartile range over 100 multi-joint trees from ArticulateArena-20K) against a constant topology penalty λ per unmatched joint ( Dλ , dashed, λ∈{0.25,0.5,1,2} ). Both schemes score joints with the same metric Eα and differ only in what an unmatched joint costs. Top: a spurious joint of growing magnitude, a missing joint whose own motion shrinks to zero, and one part split into two along the same screw. Bottom left: the range of one joint scaled by k . Bottom middle: 2,000 randomly corrupted predictions, Eαtree against D1 , colored by the first edit. Bottom right: Kendall τ rank agreement among the schemes over those predictions.
Prior per-component scores
Ours
Method
Gen ↑
TypeErr ↓
OriginErr ↓
AxisErr ↓
LimitErr ↓
SuccRate ↑
EB↓
Eαϕ↓
Eαϕ,tree↓
(%)
(%)
(m)
(rad)
(%)
(m)
Articraft
99.5
12.1
0.466
0.870
1.90 / 1.00
17.1
1.172
0.597
0.476
Articulate AnyMesh †
48.0
16.7
0.264
0.881
–
9.4
1.032
0.581
0.528
Articulate-Anything
98.5
18.3
0.604
0.629
2.19 / 0.91
15.7
1.702
0.550
0.392
ArtLLM
93.0
30.6
0.398
0.918
7.13 / 1.03
3.8
2.868
1.216
0.585
Table 4: Articulated reconstruction on 200 objects from ArticulateArena-20K . Arrows indicate the better direction; shading marks the best defined value; – denotes unavailable entries. The bootstrap intervals of Appendix D show which of these gaps the evaluation set resolves. † Articulate AnyMesh predicts no motion limits, so its E columns impute each joint’s limits from the ground truth and are excluded from the shading.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Type
Axis
Origin
Limit
Success rate
Ditto ( Jiang et al., 2022b )
✓ §
✓
✓ †
–
–
PARIS ( Liu et al., 2023a )
–
✓
✓ †
–
–
URDFormer ( Chen et al., 2024 ) ‡
–
–
–
–
–
Real2Code ( Mandi et al., 2025 )
✓
✓
✓ †
–
–
Articulate-Anything ( Le et al., 2025 )
✓
✓
✓
✓
✓
SPARK ( He et al., 2025 )
✓
✓
✓
–
–
Appendix
Table 5: Kinematic evaluation metrics reported by existing articulated-object reconstruction methods. ✓ marks a component the method reports in its own evaluation of the reconstructed kinematics, whether in its main results or in an appendix, and † marks a component reported for revolute joints only. A dash indicates that the method reports no such quantity. § Joint type accuracy is reported only in Appendix D of Jiang et al. (2022b) , on the Synthetic and Shape2Motion sets, where nearly all methods reach 100% . ‡ URDFormer infers the joint type and the link axis from the predicted mesh category and scores neither, reporting only object- and part-category accuracy, parent accuracy, a discretized part-position error, and a whole-pipeline real-robot task success rate. Geometry-only metrics such as Chamfer distance and F-score are outside the scope of this table.
Figure 6: Controlled single-parameter perturbations of ground-truth library joints, one row block per joint type; the helical row composes the prismatic + continuous encoding of a screwed bulb into one screw (pitch 5 mm per turn). Solid amber is Eα ( α=1 ; the finite → continuous column uses Eαϕ , κ=π ), dashed curves are the component scores of § 3.2 with the published thresholds τa=0.25 rad, τp=0.05 m.
Figure 7: Sensitivity to the split-norm weight α∈{0.1,…,5} , two decades around the default α=1 . Blue curves are the raw Eα , red curves the compactified Eαϕ ( κ=π ), lighter for smaller α . Rows: representative prismatic, revolute, and continuous joints. Columns: axis rotation ( 0..90∘ ), origin translation ( 0..1 m), the predicted upper limit [0,θ] , and a limit shift [l−+θ,l++θ] ; the continuous row admits only the compactified score.
Figure 8: The sweeps of Figure 7 scored under both inner products: solid brick is the kinetic-energy EB (meters), dashed blue the split-norm Eα ( α=1 , dimensionless). The two carry different units, so only shapes may be compared, not heights. The continuous row uses the compactified score, with ϕ normalized by each curve’s own norm.
Figure 9: The same representation-level errors applied to the same objects uniformly scaled by s∈[0.1,10] (log–log axes): a fixed 15∘ axis rotation, an origin offset of 0.1s m, and a halved motion range. Solid brick is EB , dashed blue Eα . The prismatic origin panel carries no curves because that error is pure gauge under both norms at every scale.
Figure 10: Sweep of the compactification scale κ . Curves show Eαϕ ( α=1 ) for κ∈{1,2,π,4,6} , lighter for smaller κ . Left three panels: a ground-truth prismatic, revolute, or helical joint of unbounded range ( l±=±∞ , the continuous case of each type) against a finite prediction [−L,L] as L grows. Right: a welded ground-truth joint against a revolute prediction whose range [0,θ] opens from zero.
Figure 11: Compactified endpoint radius tanh(∥z∥/κ) of the finite ground-truth joints of the evaluation set (per joint the larger of its two endpoints), as a CDF for each κ∈{1,2,π,4,6} , lighter for smaller κ . Rows: the norm inside ϕ (kinetic-energy, split); columns: joint type. The dotted line marks radius 0.99 , where tanh saturation would begin to collapse resolution; each panel is annotated with the κ=π statistics.
Figure 12: En compares n uniform samples of the linearized segments under the same Z2 quotient, in RMS convention so that E2=Eα/2 is the endpoint metric. Rows: prismatic and revolute joints. Columns: axis, origin, and limit sweeps. Darker curves use more samples; the dashed line is the closed-form n→∞ limit Dpath of Theorem B.17 .
Method
Eαϕ,tree
Articraft
0.476 [0.415, 0.539]
Articulate AnyMesh †
0.528 [0.436, 0.624]
Articulate-Anything
0.392 [0.337, 0.449]
ArtLLM
0.585 [0.520, 0.651]
Ditto
0.552 [0.488, 0.617]
Particulate
0.419 [0.363, 0.476]
Appendix
Table 6: The object-level score Eαϕ,tree with 95% percentile bootstrap intervals over objects, on the 200-object evaluation set. The resampling unit is the object, not the joint, and each method is resampled within its own valid outputs. A paired bootstrap over the 180 objects that the two leading methods both score puts their difference at −0.039 with a 95% interval of [−0.091,0.012] . Articulate AnyMesh † carries the ground-truth-imputed limits of Table 4 and covers 96 objects, so its interval is wider than the others.
Figure 13: Qualitative comparison on three objects from the evaluation set of Table 4 , one per row: a bottle, a box with a hinged lid, and a folding chair. The first column is the input and the remaining columns are the methods. In each panel the top render is the rest configuration and the bottom render actuates the joints, the ground-truth ones in the input column, and the kinematic tree a method returns is drawn beside its render, as in Figure 1 . A red cross marks a generation failure, so URDFormer returns no valid asset for two of the three objects. The table aggregates all valid outputs, not only these three examples.
Method
Gen
Pair
Articraft
199
158
Articulate AnyMesh †
96
79
Articulate-Anything
197
162
ArtLLM
186
105
Ditto
200
165
Particulate
182
150
Appendix
Table 7: Coverage of the 200-object evaluation in Table 4 . Gen counts objects, and Pair counts the joint pairs with finite predicted and ground-truth limits. † Counts under the ground-truth-imputed limits of Table 4 .
Method
EB (m)
EBϕ
Eαϕ
Eαϕ,tree
Articraft
1.217 [0.762, 1.927]
0.564 [0.420, 0.715]
0.702 [0.567, 0.846]
0.539 [0.423, 0.661]
Articulate AnyMesh †
1.039 [0.548, 1.836]
0.510 [0.360, 0.671]
0.590 [0.450, 0.739]
0.556 [0.436, 0.682]
Articulate-Anything
1.096 [0.549, 2.044]
0.412 [0.309, 0.527]
0.458 [0.358, 0.567]
0.303 [0.223, 0.389]
ArtLLM
3.249 [2.885, 3.616]
1.046 [0.954, 1.136]
1.183 [1.074, 1.283]
0.582 [0.468, 0.697]
Ditto
0.932 [0.512, 1.644]
0.460 [0.338, 0.590]
0.602 [0.496, 0.717]
0.544 [0.428, 0.665]
Particulate
1.176 [0.870, 1.500]
0.515 [0.409, 0.629]
0.465 [0.373, 0.556]
0.422 [0.335, 0.511]
Appendix
Table 8: Sensitivity analysis on the 67-object intersection shared by all eight methods. Entries are means with percentile 95% bootstrap intervals in brackets. EB is computed on the pairs with finite predicted and ground-truth limits; the other three scores are dimensionless and defined on all 67 objects. † Articulate AnyMesh predicts no motion limits, so its limits are imputed from the ground truth as in Table 4 and its row reads as a geometry-only diagnostic rather than a competitive score.
Figure 14: Category sizes of ArticulateArena-20K by rank (bars, log scale) and the cumulative share of objects (curve). The largest 77 categories hold half of the objects and the largest 204 hold 80%.
Department of Engineering, University of Cambridge, Cambridge CB2 1PZ, U.K. · School of Engineering and Design, Technical University of Munich · Munich Center for Machine Learning (MCML), Munich, Germany.