Generalization in robotic manipulation requires policies to perform tasks across diverse unseen object instances that vary in shape, size, and pose. However, conventional behavior cloning (BC) methods often overfit to instance-specific geometry and appearance, limiting transfer to novel objects. We introduce KeyGen, a framework that learns canonicalized semantic 3D keypoints from point clouds and uses them as structured object-centric representations for policy learning. A visuomotor diffusion policy conditions on these keypoints together with object-centric geometry to predict full manipulation trajectories, enabling consistent geometric correspondence across object instances. To evaluate category-level generalization, we construct a photorealistic simulation benchmark with three manipulation tasks and a planning-driven data generation pipeline that produces expert trajectories across diverse object instances. Experiments show that KeyGen significantly outperforms prior methods on both seen and unseen objects under pose variation, scales effectively with additional demonstrations per object, maintains robustness to object rescaling, and achieves strong performance in both simulation and real-world manipulation.
Figures & tables
Figure 1 : We propose KeyGen , a framework that enables generalization across unseen object instances by leveraging semantically aligned 3D keypoints. KeyGen uses unsupervised, task-agnostic keypoints as structured object representations, allowing policies to reason about geometry and pose for precise and transferable manipulation.
Figure 2 : Overview of KeyGen . KeyGen segments multi-view RGBD images to create an object-centric point cloud, which is used to extract consistent semantic 3D keypoints. Together with proprioception and noisy action, the encoded object-centric point cloud and keypoints serve as conditioning features in the Diffusion Process. At inference time, the action is iteratively denoised to predict robot trajectory.
Figure 3 : Category-specific canonical alignment and 3D keypoint detection. During training, two random partial scans of the same object are aligned to a shared canonical frame by a pretrained, category-specific Canonical Form Matching network and then processed by the keypoint detector F ; separation ( Lsep ), heuristic ( Lheur ), and consistency ( Lcons ) losses enforce landmark spread, saliency, and cross-view repeatability. At inference, a single scan is aligned to canonical space, F predicts keypoints, and they are reprojected to the original view via the estimated pose.
Figure 4 : Left : Representative rollout examples of our method on the three manipulation tasks. Right : Common failure modes observed in baseline methods, including grasp misalignment, incorrect placement, and orientation errors.
Method
Pouring
Collect Knife
Stacking
Easy
Hard
Easy
Hard
(no pose var.)
Seen Instances (in-distribution)
ACT [ 35 ]
85%
35%
95%
35%
30%
DP3 [ 36 ]
80%
30%
90%
20%
20%
3DDA [ 37 ]
95%
40%
95%
45%
65%
BAKU [ 6 ]
90%
55%
80%
45%
60%
Table I: Category-level generalization across tasks. Success rates (%) on both seen instances (top block) and unseen instances (bottom block). Each policy is trained on a single object and evaluated with 20 rollouts per configuration. Easy : random position; Hard : random position + orientation. Stacking involves negligible pose variation.
Variant
Pouring
Collect Knife
Stacking
Seen
Unseen
Seen
Unseen
Seen
Unseen
KeyGen w/o Object-centric
55%
29.6%
45%
32.5%
60%
44.2%
KeyGen w/o Keypoints
90%
45%
90%
59.2%
95%
67.5%
KeyGen with 2D Keypoints
65%
10.4%
30%
22.5%
90%
60%
KeyGen (Full)
90%
56.3%
85%
71.7%
85%
73.8%
Table II: Ablation study : success rates (%) on seen and unseen objects across three tasks. We evaluate removing object-centric cropping, omitting keypoints, and replacing 3D keypoints with 2D image-based keypoints.
Figure 5 : Robustness to object-scale shifts across three tasks. Policies are trained at scale 1.0 and evaluated zero-shot at isotropic multipliers {0.5, 0.7, 0.9, 1.0, 1.2} (x-axes). Each marker shows success rate (%) for a method (rows); color encodes success-rate bins (legend). KeyGen stays high under moderate rescaling and degrades mainly at the extremes, while baselines remain low or fluctuate.
Figure 6 : Data efficiency with more demonstrations per object. We fix the training object set and vary trajectories per object (10, 20, 50, 100). Curves show success on held-out objects for two tasks ( pouring , collect knife ). KeyGen converts additional trajectories into steady gains and continues to improve at the highest data regime, while ACT grows slowly after early gains, 3DDA improves then flattens, and DP3 changes little.
Figure 7 : Real-world qualitative results. Left: execution snapshots for two real-robot tasks— pouring (top) and align shoe (bottom)—illustrating the task completion process. Right: rollouts on unseen object instances, demonstrating that KeyGen generalizes across different cups and shoes with varying geometry and appearance.
Figure 8 : Real-world performance on seen vs. unseen objects. Success rates for two real-robot tasks— align shoe (top) and pouring (bottom)—evaluated on training instances (Seen, left) and held-out object instances (Unseen, right). KeyGen achieves the highest success in both tasks and settings, outperforming ACT, DP3, and GenDP, with the largest gains on unseen objects.
Manipulation policies deployed in uncontrolled real-world scenarios are faced with great in-category geometric diversity of everyday objects. In order to function robustly under such variations, policies need to work in a category-level manner, i.e. knowing how to interact with any object in a certain category, instead of only a specific one seen during training. This in-category generalizability is usually nurtured with shape-diversified training data; however, manually collecting such a corpus of data is infeasible due to the requirement of intense human labor and large collections of divergent objects at hand. In this paper, we propose ShapeGen, a data generation method that aims at generating shape-variated manipulation data in a simulator-free and 3D manner. ShapeGen decomposes the process into two stages: Shape Library curation and Function-Aware Generation. In the first stage, we train spatial warpings between shapes mapping points to points that correspond functionally, and aggregate 3D models along with the warpings into a plug-and-play Shape Library. In the second stage, we design a pipeline that, leveraging established Libraries, requires only minimal human annotation to generate physically plausible and functionally correct novel demonstrations. Experiments in the real world demonstrate the effectiveness of ShapeGen to boost policies' in-category shape generalizability. Project page: https://wangyr22.github.io/ShapeGen/.
RGB-based imitation learning requires many demonstrations to generalize to unseen objects or scenes, motivating research into intermediate representations to improve generalization for robotic manipulation. Visual foundation models enable one-shot extraction of keypoints to provide such representation. However, it remains unclear how to integrate them into imitation learning optimally and when they outperform alternative representations. We combine approaches from previous works on keypoint imitation learning (KIL) and investigate several design choices to provide practical guidelines. Using over 2000 real-world rollouts, we also assess the generalization capabilities of KIL to unseen objects and scene variations. KIL achieves a 75% overall success rate across five tasks, significantly outperforming the RGB baseline (47%) and performing on par with S2-diffusion (73%). Finally, we explore the limitations of the foundation models used for keypoint extraction and extend KIL to tasks with multiple object instances. Our results confirm KIL as a data-efficient approach for robot learning, though it does not outperform alternative representations and inherits limitations of the foundation models used for keypoint extraction. All rollout videos, demonstrations, and results are available at https://kil-manipulation.github.io/.
Thomas Lips, Marco Moletta, Michael C. Welle +2
AI and Robotics Lab, IDLAB-AIRO, Ghent University-imec · Robotics, Perception and Learning Lab, (RPL), EECS, KTH Royal Institute of Technology. · INCAR Robotics AB, Stockholm, Sweden.
Keypose-based manipulation decomposes tasks into critical waypoints to simplify policy learning for long-horizon tasks, but existing approaches rely on task-specific heuristics or manual annotation to extract keyposes from demonstrations. We present an automatic trajectory labelling pipeline for grasp-related tasks. This pipeline combines vision-language models (VLMs) for semantic event detection with classical trajectory analysis for precise temporal alignment, requiring VLM inference only on one single demo among repeating ones per task. Using the labelled data, we train a keypose-guided Diffusion Policy (DP) that exploits keypose conditioning to intervene demonstration distributions. We explore the possibility to apply this property for cross-embodiment transfer: candidate keyposes are sampled and filtered via a reachability map, steering the policy toward kinematically feasible keyposes for the target robot. As a preliminary feasibility study, experiments on two robomimic tasks show that the labelled data produces policies matching a standard DP baseline, and that reachability-filtered keypose conditioning may benefit zero-shot transfer on the multimodal insertion task when feasible candidates are available.
Yupu Lu, Hang Xu, Yizhou Chen +1
School of Computing and Data Science, The University of Hong Kong, HKSAR.