Enabling robots to estimate the kinematic parameters of articulated objects unlocks a wide range of capabilities for interaction and manipulation. The estimation has to happen from the information the robot currently observes, often just a single RGB image of an object it has never seen before. Current single-image approaches couple articulation part segmentation with articulation estimation, making their predictions vulnerable to missed detections and incorrect part associations, and they regress metric 3D geometry that a single view fixes only up to scale. We present QueryArt, a model that estimates articulation parameters from a single RGB image, a 2D query point, and camera intrinsics. QueryArt is trained to estimate the 3D articulation geometry relative to the queried point and in units of its depth, which keeps its target identifiable from the image alone. A single depth measurement at the query point then supplies the scale and recovers the metric parameters. We train QueryArt on a curated mixture of synthetic and real-world articulation datasets. We evaluate QueryArt on several benchmarks and compare it against recent baselines. QueryArt outperforms recent baselines on most articulation metrics, including on out-of-distribution data. To demonstrate the model's capabilities in real-world settings, we evaluate QueryArt on a mobile manipulator across 57 manipulation trials spanning 16 object parts and five viewpoint classes, achieving a 70.2% success rate. We provide code and videos at: https://abwerby.github.io/queryart/
Figures & tables
Fig. 1: QueryArt estimates articulation in a normalized coordinate space from a single RGB image and a 2D query point. Depth at the query point enables lifting the prediction to metric 3D space. The magenta arrow illustrates motion consistent with the predicted articulation. Bottom: real-world deployment on a mobile manipulator.
Fig. 2: QueryArt architecture. A frozen vision transformer (DINOv3 ViT-B/16) encodes the RGB image; early and final-block features are fused into a three-scale pyramid Fimg . The query embedding Fq is built from the same point-wise visual features sampled at the query, a Fourier encoding of the pixel location, and a hand-crafted calibration vector derived from the intrinsics K . Fq is added to two randomly initialized learned tokens: Type and Line , which then pass through two transformer blocks that cross-attend to Fimg and self-attend to each other. The prediction heads output the joint classification and the axis parameters for both moving types. Right: the predicted geometry in query-normalized coordinates. The offset r^ is orthogonal to a^rev , so it points from the unprojected query p~ to the nearest point o~=p~+r^ on the revolute axis. Expressing r^ in units of the query depth makes the axis recoverable up to a single global scale factor.
Axis error ↓
Type recall ↑
dh↓
RMSE ↓
S↑
Method
R
P
R
P
(m)
R
P
(%)
OPDFormer-C [ 4 ]
16.07
22.18
61.11
4.76
0.874
0.315
0.083
6.30
MOPD [ 5 ]
11.94
—
87.37
0.00
0.823
0.281
—
11.53
A3VLM [ 6 ]
36.34
77.26
7.85
0.12
0.768
0.318
0.270
0.00
3DOI [ 16 ]
5.13
3.30
67.01
86.05
0.297
0.091
0.012
48.54
VidBot [ 7 ]
—
—
—
—
—
0.302 †
0.304 †
—
TABLE I: Held-out split results on four datasets. Articulation-macro means grouped by GT joint type. R/P : revolute/prismatic; axis error in degrees, recall in %, dh and RMSE in meters. † : native trajectory protocol.
Fig. 3: Qualitative articulation predictions on objects unseen during training: Arti4D [ 10 ] (left), self-captured scenes (middle), HOI! [ 41 ] (right). Each scene is queried at several points. Top: input image with the query points and the projected motion trajectory. Bottom: the same predictions on the point cloud, with the predicted axis of motion and the 3D trajectory.
Axis error ↓
Type recall ↑
dh↓
RMSE ↓
S↑
Method
R
P
R
P
(m)
R
P
(%)
OPDFormer-C [ 4 ]
21.69
—
93.79
0.00
0.471
0.170
—
15.14
MOPD [ 5 ]
15.90
6.11
93.75
2.59
0.545
0.178
0.023
18.70
A3VLM [ 6 ]
51.76
88.13
4.44
0.41
0.570
0.231
0.300
0.00
3DOI [ 16 ]
9.64
6.90
69.50
76.56
0.241
0.090
0.025
58.58
VidBot [ 7 ]
—
—
—
—
—
0.125 †
0.074 †
—
TABLE II: Out-of-distribution results on HOI! [ 41 ] and Arti4D [ 10 ] , both excluded from training. Columns and symbols as in Tab. I .
Fig. 4: Real-robot trials by viewpoint. Each trial is assigned to one of five viewpoint classes and flows to a success or a failure outcome.
Axis error ↓
Type recall ↑
dh↓
RMSE ↓
S↑
Variant
R
P
R
P
(m)
R
P
(%)
Articulate3D
Full
6.12
9.38
96.20
89.25
0.201
0.062
0.035
70.03
w/o pyramid
+2.11
-1.49
+2.27
+0.24
+0.0270
+0.0074
-0.0056
-1.51
dense attention
-0.55
-0.07
+2.14
-1.95
+0.0603
+0.0137
-0.0003
-6.48
w/o calibration
+0.41
+0.46
-0.35
-1.71
+0.0177
+0.0045
+0.0017
-2.73
TABLE III: Ablations on Articulate3D and HOI! [ 41 ] . Full: absolute values; other rows: signed change from Full, recall in percentage points. Units as in Tab. I . Green : improvement; red : regression.
Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. Existing approaches often rely on prior knowledge about the objects, such as the number or type of joints. Some of these approaches also fail to recover occluded joints that are only revealed during interaction. Others require large numbers of multi-view images for every object, which is impractical in real-world settings. Furthermore, prior works neglect the order of manipulations, which is essential for many multi-DoF objects where one joint must be operated before another, such as a dishwasher. We introduce PokeNet, an end-to-end framework that estimates articulation models from a single human demonstration without prior object knowledge. Given a sequence of point cloud observations of a human manipulating an unknown object, PokeNet predicts joint parameters, infers manipulation order, and tracks joint states over time. PokeNet outperforms existing state-of-the-art methods, improving joint axis and state estimation accuracy by an average of over 27% across diverse objects, including novel and unseen categories. We demonstrate these gains in both simulation and real-world environments.
Anmol Gupta, Weiwei Gu, Omkar Patil +2
School of Computation and AI, ASU, Tempe · AI Institute, Seoul National University, Seoul, South Korea
Understanding articulated objects is fundamental for robotic interaction, requiring accurate rigid-part discovery and the recovery of their kinematic relations. Existing approaches often treat articulation as a by-product of reconstructed geometry or recover it through per-instance optimization. We instead build on the hypothesis that articulation is directly observable from persistent motion: points on the same rigid part move coherently, while relative motion between parts reveals their kinematic constraints. We present Track2Art, a motion-centric framework for recovering structured articulated objects from RGB-D interaction videos. Track2Art lifts tracked image points into persistent 3D trajectories and combines pretrained tracking features, visual descriptors, and explicit trajectory geometry. These representations are grouped into a variable number of rigid-part hypotheses and subsequently used to recover directed kinematic relations, joint types, and joint geometry through rotation-equivariant learned--analytic reasoning. On the aligned 20-object PartNet-Mobility suite, Track2Art achieves 0.695 Point IoU and 0.410 end-to-end J@20, while requiring neither ground-truth part counts nor test-time optimization.
Xiaotong Li, Yixiong Jing, Junsheng Ding +4
Department of Engineering, University of Cambridge, Cambridge CB2 1PZ, U.K. · School of Engineering and Design, Technical University of Munich · Munich Center for Machine Learning (MCML), Munich, Germany.
The kinematics of an articulated object is often ambiguous from vision alone. Interaction resolves the ambiguity, and active perception methods exploit this by searching for the single action that most sharpens a belief over the kinematic parameters at each step. Such greedy search cannot be extended over a horizon without forward models of the contact and inertial dynamics, which are themselves unknown. We instead amortize action selection into training. We maintain a belief distribution over joint type and parameters, initialized from a generative prior and updated by Bayesian filtering on the observed part motion. To condition the policy on this belief, we render it as a per-point articulation flow field, the motion that the current posterior predicts for every point on the object. Carrying the inductive bias of articulated motion, this representation generalizes better than a latent encoding of the belief or flow tracked from observation. We train the policy with reinforcement learning, rewarding the entropy that each interaction removes from the posterior, so that informative exploration becomes learned behavior rather than a search at every step. Our method outperforms previous approaches across door and drawer manipulation on the PartManip benchmark, and reaches 61.7% success on ArticuRiddle, a new dataset of objects whose appearance implies the wrong articulation, against 44.4% for the best previous method. Project Website: https://hiddenkinematics.github.io/
Ruiyao Liu, Boshu Lei, Zhuoyang Pan +1
GRASP Lab, University of Pennsylvania, Philadelphia, PA 19104, USA