The kinematics of an articulated object is often ambiguous from vision alone. Interaction resolves the ambiguity, and active perception methods exploit this by searching for the single action that most sharpens a belief over the kinematic parameters at each step. Such greedy search cannot be extended over a horizon without forward models of the contact and inertial dynamics, which are themselves unknown. We instead amortize action selection into training. We maintain a belief distribution over joint type and parameters, initialized from a generative prior and updated by Bayesian filtering on the observed part motion. To condition the policy on this belief, we render it as a per-point articulation flow field, the motion that the current posterior predicts for every point on the object. Carrying the inductive bias of articulated motion, this representation generalizes better than a latent encoding of the belief or flow tracked from observation. We train the policy with reinforcement learning, rewarding the entropy that each interaction removes from the posterior, so that informative exploration becomes learned behavior rather than a search at every step. Our method outperforms previous approaches across door and drawer manipulation on the PartManip benchmark, and reaches 61.7% success on ArticuRiddle, a new dataset of objects whose appearance implies the wrong articulation, against 44.4% for the best previous method. Project Website: https://hiddenkinematics.github.io/
Figures & tables
Fig. 1: Visual appearance can be misleading. We change a drawer’s joint from prismatic to revolute without altering its appearance. Both methods open the original drawer, but on the modified one the baseline is misled by the geometry, whereas ours refines its belief through interaction and adapts.
Fig. 2: Pipeline Overview. Given the initial observation, we use a generative prior to initialize a set of particles, where each particle represents a hypothesis of the object’s articulation model. At each interaction step, each hypothesis is converted into a per-point flow field. These flow fields are weighted-averaged over the particles. The flow field is first concatenated with the current point cloud as Ot′ and then fused with the encoded robot’s proprioceptive state as input to the actor. The predicted action is executed in the environment, and the observed motion is used to update the particle weights through a Bayes filter. The information reward, defined as the reduction in belief entropy from before to after the update, is combined with the task reward to train the policy using PPO.
Fig. 3: Flow field updates during interaction at time t . As the robot interacts with the object, hypotheses inconsistent with the observed motion are suppressed and the posterior gradually concentrates. The flow is updated accordingly, providing the policy with increasingly accurate motion cues.
Fig. 4: Overview of ArticuRiddle. We construct challenging objects through (a) handle-pose edits and (b) joint-parameter edits that preserve appearance, both testing manipulation under ambiguous geometric cues.
Category
Method
Train
Val
ArticuRiddle
Acc. ↑
ASR 10↑
ASR 20↑
Acc. ↑
ASR 10↑
ASR 20↑
Acc. ↑
ASR 10↑
ASR 20↑
Door
AKM [ 4 ]
79.0 ± 1.9
48.2 ± 1.3
55.9 ± 2.5
84.1 ± 1.6
29.6 ± 5.6
33.3 ± 5.9
64.0 ± 3.5
30.7 ± 6.1
32.7 ± 5.8
H-SAUR [ 3 ]
81.7 ± 1.9
40.3 ± 3.6
46.6 ± 4.6
85.2 ± 2.7
35.1 ± 5.4
43.2 ± 8.1
68.1 ± 1.9
39.0 ± 2.9
50.3 ± 2.9
Ours
82.7 ± 1.3
62.4 ± 1.3
72.7 ± 2.1
87.8 ± 2.2
59.2 ± 5.1
70.1 ± 3.0
78.4 ± 5.2
57.7 ± 5.6
65.4 ± 3.3
Drawer
AKM [ 4 ]
75.9 ± 1.6
42.1 ± 1.9
49.7 ± 3.8
79.2 ± 4.4
41.5 ± 4.0
52.5 ± 6.5
54.1 ± 4.7
35.8 ± 2.2
41.7 ± 1.8
H-SAUR [ 3 ]
73.0 ± 2.0
43.0 ± 2.3
49.7 ± 0.4
73.5 ± 1.9
38.1 ± 0.3
54.3 ± 7.1
59.0 ± 1.3
29.0 ± 5.5
36.2 ± 3.3
TABLE I: Articulation estimation on PartManip and ArticuRiddle.
Method
Door
Drawer
ArticuRiddle
Train
Val
Train
Val
H-SAUR [ 3 ]
21.8 ± 1.2
27.0 ± 1.3
33.5 ± 1.8
32.7 ± 1.5
44.4 ± 2.2
AKM [ 4 ]
25.3 ± 2.0
19.0 ± 1.6
26.7 ± 1.1
22.6 ± 1.4
27.0 ± 2.5
Wang et al. [ 35 ]
37.4 ± 2.5
39.7 ± 5.4
44.2 ± 3.3
27.8 ± 10.6
18.2 ± 3.9
PartManip [ 10 ]
44.5 ± 0.7
40.6 ± 2.7
67.1 ± 2.6
74.3 ± 4.3
32.4 ± 1.4
Obs-flow
59.4 ± 1.9
43.6 ± 3.6
76.7 ± 1.8
88.0 ± 1.1
53.8 ± 3.4
TABLE II: Manipulation success rate.
Fig. 5: Reward alignment and opening progress on the training set. Top : episodes collected with our trained policies are scored by three different rewards; bars give the fraction with a final axis error above 10∘ in the lowest- (hatched) and highest-scoring (solid) quarter. The results suggest optimizing the entropy-reduction return may lead to more accurate estimates. Bottom : median opening over assets for the policy trained with each reward; the dashed line marks the success threshold. Error bars and bands are bootstrapped 95% confidence intervals, over episodes (top) and over assets (bottom).
Variant
Door
Drawer
ArticuRiddle
Train
Val
Train
Val
w/o RL
57.6 ± 1.2
40.2 ± 2.4
82.7 ± 3.0
87.7 ± 1.0
48.1 ± 1.1
δq
63.9 ± 1.5
42.9 ± 4.2
84.4 ± 1.0
90.4 ± 2.3
55.2 ± 2.8
DKL
64.4 ± 1.1
42.9 ± 3.2
87.1 ± 3.9
93.9 ± 1.3
58.6 ± 4.6
ΔH (Ours)
68.2 ± 1.3
51.3 ± 1.9
89.0 ± 3.0
95.4 ± 1.5
61.7 ± 1.1
TABLE III: Reward ablation on manipulation success rate.
Fig. 6: Real-world manipulation examples. Each row shows the object followed by point-cloud visualizations at successive interaction steps. The rendered flow field shows how observed part motion drives the articulation belief to converge toward the true joint.
Object
Axis ( ∘ ) ↓
Pivot (cm) ↓
SR (%) ↑
Door 1
2.5
2.2
60.0
Door 2
3.1
2.6
70.0
Drawer 1
1.7
–
90.0
Drawer 2
3.6
–
80.0
TABLE IV: Real-world articulation estimation errors and manipulation success rates. A trial succeeds when the door opens by more than 30∘ or the drawer by more than 20% of its opening range.
Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects. By embedding interpretable geometric and kinematic priors into policy learning, KAI provides a strong inductive bias aligned with the underlying structure of articulated motion. This design effectively improves sample efficiency, with gains particularly pronounced in low-data regimes: across six simulation tasks, our method achieves an average success rate of 82.9%, matching or surpassing baseline performance while using only half the demonstration data. Our method also exhibits robust generalization to unseen backgrounds and visual distractors, transferring from a single clean training environment to cluttered real-world scenes. KAI's action-agnostic design further enables co-training with human interaction videos to enhance real-world robustness: under diverse visual distractions, our method with video co-training achieves over 70% average success rate.
Yaping Li, Zhaxizhuoma, Qiaojun Yu +3
The Chinese University of Hong Kong · Shanghai AI Laboratory · Shanghai Jiao Tong University
Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. Existing approaches often rely on prior knowledge about the objects, such as the number or type of joints. Some of these approaches also fail to recover occluded joints that are only revealed during interaction. Others require large numbers of multi-view images for every object, which is impractical in real-world settings. Furthermore, prior works neglect the order of manipulations, which is essential for many multi-DoF objects where one joint must be operated before another, such as a dishwasher. We introduce PokeNet, an end-to-end framework that estimates articulation models from a single human demonstration without prior object knowledge. Given a sequence of point cloud observations of a human manipulating an unknown object, PokeNet predicts joint parameters, infers manipulation order, and tracks joint states over time. PokeNet outperforms existing state-of-the-art methods, improving joint axis and state estimation accuracy by an average of over 27% across diverse objects, including novel and unseen categories. We demonstrate these gains in both simulation and real-world environments.
Anmol Gupta, Weiwei Gu, Omkar Patil +2
School of Computation and AI, ASU, Tempe · AI Institute, Seoul National University, Seoul, South Korea
Enabling robots to estimate the kinematic parameters of articulated objects unlocks a wide range of capabilities for interaction and manipulation. The estimation has to happen from the information the robot currently observes, often just a single RGB image of an object it has never seen before. Current single-image approaches couple articulation part segmentation with articulation estimation, making their predictions vulnerable to missed detections and incorrect part associations, and they regress metric 3D geometry that a single view fixes only up to scale. We present QueryArt, a model that estimates articulation parameters from a single RGB image, a 2D query point, and camera intrinsics. QueryArt is trained to estimate the 3D articulation geometry relative to the queried point and in units of its depth, which keeps its target identifiable from the image alone. A single depth measurement at the query point then supplies the scale and recovers the metric parameters. We train QueryArt on a curated mixture of synthetic and real-world articulation datasets. We evaluate QueryArt on several benchmarks and compare it against recent baselines. QueryArt outperforms recent baselines on most articulation metrics, including on out-of-distribution data. To demonstrate the model's capabilities in real-world settings, we evaluate QueryArt on a mobile manipulator across 57 manipulation trials spanning 16 object parts and five viewpoint classes, achieving a 70.2% success rate. We provide code and videos at: https://abwerby.github.io/queryart/
Abdelrhman Werby, Fabio Scaparro, Kai O. Arra
Socially Intelligent Robotics Lab, Institute for Artificial Intelligence, University of Stuttgart, Germany