DROM: A Language-Guided Diffusion Framework for Multi-Skill Robotic Manipulation
Authors: Vincenzo Pomponi, Rocco Felici, Paolo Franceschi, Stefano Baraldo, Oliver Avram, Loris Roveda, Luca Maria Gambardella, Anna Valente
Organizations: Institute of Systems and Technologies for Sustainable Production (ISTePS), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland. · Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland. · Mechanical Department, Politecnico di Milano (PoliMi), Via Giuseppe Candiani, Milan, 20155, Italy. · Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Faculty of Informatics, Universit`a della Svizzera Italiana (USI), Via Buffi 13, Lugano, CH-6900, Ticino, Switzerland.
Learning robust manipulation policies for diverse, long-horizon tasks from limited demonstrations remains a fundamental challenge in robotics. We present DROM, a language-guided diffusion framework that enables robots to learn, represent, and compose multiple manipulation skills within a single generative policy. DROM leverages Dynamic Movement Primitives (DMPs) to augment a small set of expert demonstrations into expressive multi-skill datasets, substantially reducing data collection while improving spatial generalization beyond the demonstrated workspace. Building upon Motion Planning Diffusion (MPD), we extend the diffusion architecture to support language-conditioned multi-skill trajectory generation through cross-attention, allowing a single model to generate skill-consistent motions for a diverse set of manipulation primitives, including orientation-sensitive behaviors that are difficult to design using conventional motion planning or hard-coded controllers. For long-horizon manipulation, a large language model decomposes high-level operator requests into executable sequences of skills, enabling natural language interaction and autonomous task execution. We validate DROM on a Franka Emika Panda robot, a FANUC CRX25ia robot, and in MuJoCo simulation across a wide range of manipulation tasks. Experimental results demonstrate that DROM outperforms Motion Planning Diffusion and Behavior Cloning baselines, achieves robust multi-skill generalization, and composes learned skills to reliably execute long-horizon manipulation tasks from natural language instructions using only a limited number of human demonstrations. Datasets, simulation environments, and more at https://github.com/automation-robotics-machines/drom.
Figures & tables
Figure 1 : Illustration of the two operational modalities of DROM: (a) language-guided single-skill execution, where an operator query is embedded and mapped to a manipulation primitive to generate a skill-consistent trajectory; and (b) language-guided long-horizon planning, where a language model decomposes a high-level objective into an ordered sequence of skills, each executed via diffusion-based trajectory generation.
Figure 2 : Overview of the diffusion-based training pipeline of DROM: (A) collection of a single expert demonstration per skill; (B) synthetic dataset expansion using Dynamic Movement Primitives (DMPs); (C) training of the skill-conditioned diffusion model on the augmented multi-skill dataset; and (D) deployment in simulation and on real robotic platforms.
Figure 3 : The set of manipulation skills considered in this study for the FANUC CRX25ia robot (a–i) and the Franka Emika Panda (j–o).
Skill
Demos
Skill
Demos
Co-Transport
1
Pick Battery
2
Shelf
1
Place Battery
2
Slide
1
Pick Bottle
4
Stir
1
Place Bottle
4
Pick Cube
1
Pick Cup
4
Stack Cube
1
Pour Liquid
4
Table 1 : Number of demonstrations collected for each manipulation skill.
Figure 4 : Architecture of the proposed skill-conditioned diffusion model. At each reverse diffusion step, a Temporal U-Net denoiser predicts the noise of the trajectory conditioned on the diffusion timestep, the environment state, and the language embedding of the user prompt. Cross-attention layers at every level of the U-Net allow the model to attend to skill-specific semantic information while preserving temporal structure through the U-Net backbone.
Skill
1 Demo
2 Demos
4 Demos
Pick Battery
48%
92%
92%
Place Battery
50%
98%
98%
Pick Bottle
24%
72%
80%
Place Bottle
24%
78%
84%
Pick Cup
28%
46%
86%
Pour Liquid
22%
56%
92%
Table 2 : Effect of the number of demonstrations on Diffusion Model performance. Success rates obtained for orientation-sensitive manipulation skills when increasing the number of demonstrations provided to the DMP augmentation process.
Object
Px
Py
Pz
Rz
Battery
23%
21%
7%
16%
Bottle
24%
24%
8%
9%
Button
24%
24%
8%
5%
Co-Transport Object
28%
22%
7%
23%
Pouring Container
31%
19%
3%
19%
Shelf
17%
29%
2%
33%
Table 3 : Workspace of each manipulated object, reported in terms of translational and rotational variations.
FANUC CRX25ia
Franka Research 3
Skill
MPD
DROM
Skill
MPD
DROM
Pick Battery
76%
98%
Pick Cube
78%
90%
Place Battery
82%
96%
Stack Cube
44%
94%
Pick Bottle
60%
84%
Grab Handle
38%
90%
Place Bottle
76%
94%
Open Drawer
50%
86%
Pick Cup
38%
86%
Place in Drawer
62%
92%
Table 4 : Comparison between DROM and MPD on the considered manipulation skills. The left half reports results on the FANUC CRX25ia, while the right half reports results on the Franka Emika Panda. DROM consistently outperforms MPD across both robotic platforms by learning multiple manipulation skills within a single language-conditioned diffusion model.
Figure 5 : Simulation environments. The two environments are designed to closely replicate the manipulation setups executed on the real Franka arm, while introducing structured, multi-stage objectives that decompose naturally into sequential subtasks.
Env.
Skill
MPD
BC
DROM
MugCleanup
Pick Cube
59.33±2.49
96.00±2.00
97.33±1.15
Stack Cube
60.67±0.94
92.67±2.31
94.67±0.94
Stack
Grab Handle
28.00±1.63
92.00±3.46
96.00±1.63
Open Drawer
58.00±1.63
90.67±0.94
96.67±0.94
Pick Mug
28.00±2.83
85.33±4.11
88.00±1.63
Place in Drawer
30.67±0.94
92.00±1.63
94.67±2.49
Table 5 : Comparison of DROM, MPD, and BC, in the MuJoCo simulation environments: Stack and Mug Cleanup.
Figure 6 : Qualitative examples of multimodal trajectory generation. (a) Pick Bottle : depending on the bottle orientation, the diffusion model generates different approach trajectories to insert a gripper finger into the handle. (b) Grab Handle : the model adapts the grasping motion to two drawers with different handle geometries, producing the corresponding manipulation strategy.
Task
# Skills
# Objects
SR
Pick and Place
2
2
86.0%
Pour
2
1
70.0%
Stack
2
2
67.5%
Hide
5
1
70.4%
Clean Up
5 – 9
1 – 3
60.9%
Table 6 : End-to-end evaluation of DROM. Success rate (SR) achieved on long-horizon manipulation tasks of increasing complexity.
Task
Skill Composition
Sequence Execution
Overall
Pick and Place
100%
86.0%
86.0%
Pour
87.5%
80.0%
70.0%
Stack
75.0%
90.0%
67.5%
Hide
80.0%
88.0%
70.4%
Clean Up
72.5%
84.0%
60.9%
Table 7 : Success rates of skill composition, sequence execution and overall for Stack and Clean Up tasks.
Category
Failure Mode
#
%
Language Model
Incorrect skill sequence
42/250
16.8
Diffusion Model
Failed grasp
11/250
4.4
Placement failure
7/250
2.8
Collision
7/250
2.8
Unfeasible trajectory
1/250
0.4
Goal not reached
10/250
4
Table 8 : Failure analysis of DROM. Failures are grouped according to the corresponding module of the framework.
Skill
Precision
Recall
F1-score
SR
Pick Cube
96.5%
83.0%
89.2%
89.0%
Stack Cube
97.6%
82.0%
89.1%
85.0%
Grab Handle
74.6%
85.0%
79.4%
82.0%
Open Drawer
85.6%
89.0%
87.3%
85.0%
Place in Drawer
75.0%
93.0%
83.0%
93.0%
Close Drawer
96.6%
85.0%
90.4%
83.0%
Table 9 : Language encoder evaluation on a balanced test set of 600 natural language instructions. Precision, recall, F1-score, and class-wise success rate (SR) quantify the encoder’s ability to associate operator requests with the corresponding manipulation primitive.
Figure 7 : Confusion matrix of the language encoder on the six manipulation primitives executed with the Franka Emika Panda robot.
Skill
Language Recognition
Imitation Learning
Overall Success
Pick Cube
89.0%
90.0%
80.1%
Stack Cube
85.0%
94.0%
79.9%
Grab Handle
82.0%
90.0%
73.8%
Open Drawer
85.0%
86.0%
73.1%
Place in Drawer
93.0%
92.0%
85.6%
Close Drawer
83.0%
88.0%
73.0%
Table 10 : Evaluation of language-guided diffusion planning , in terms of language recognition accuracy, skill execution success rate, and request-to-goal success rate.
Scaling imitation learning to diverse multi-task robot manipulation remains challenging due to suboptimal demonstrations, behavioral multi-modality, and destructive interference across tasks. While skill-based methods offer a promising direction by decomposing behaviors into reusable abstractions, existing approaches often learn skills that are either biased toward linguistic structure or lack semantic alignment across tasks, limiting generalization. In this work, we propose AtomSkill, a novel framework that learns a semantically aligned Atomic Skill Space from demonstrations and enables robust long-horizon execution through keypose imagination. Our method introduces: (1) semantic contrastive skill alignment, which partitions demonstrations into variable-length atomic skills and employs a contrastive objective to jointly enforce semantic consistency and temporal coherence, yielding a compact and reusable skill library; and (2) action decoding with keypose imagining, where the policy predicts both a skill's terminal keypose and immediate actions, thereby supporting progress-aware skill transitions. During inference, an atomic skill diffusion sampler generates plausible skill sequences, while predicted keyposes autonomously trigger smooth skill chaining. Extensive experiments in simulation and real-world settings show that AtomSkill consistently outperforms state-of-the-art imitation learning and skill-based baselines. Project page: https://atom-skill.github.io.
Yihang Zhu, Weiqing Wang, Shijie Wu +2
ShanghaiTech University, Shanghai, China · InstAdapt
Successfully automating dexterous, long-horizon robotic manipulation requires frameworks capable of both high-level reasoning and fine-grained execution. Traditional task and motion planning (TAMP), while excellent at symbolic planning, is often brittle in contact-rich operations. Simultaneously, imitation learning (IL), while effective in manipulation tasks with visual feedback, is limited by its low capability in spatial generalization and multi-stage operation. To reconcile their complementary strengths and limitations, we propose DR-LfD (Decomposed and Reorganized Skills Learned from Demonstrations), a framework that seamlessly integrates visuomotor policies into a TAMP-gated decision-making system. Based on contact relationships, DR-LfD decomposes human demonstrations into atomic skills, which are reproduced as visuomotor policies or object-centric primitives. The initiation, termination, and constraints of the visuomotor policies are carefully modeled and implemented in a TAMP-compatible form, enabling reorganization of skills learned from different sources. DR-LfD transforms the learning problem from one requiring exponential demonstration data over possible skill sequences to one whose demonstration burden scales with the number of distinct skill types, with limited data for each skill. Through comprehensive real-world and simulation benchmarking across diverse scenarios, we demonstrate the strong performance of DR-LfD on tasks involving multiple steps, unseen setups, and physical constraints. Project website: https://dr-lfd.github.io/DR-LfD-website.
Yizhou Chen, Hang Xu, Dongjie Yu +7
The University of Hong Kong, Hong Kong · JD.com (JingDong) · Nanyang Technological University (NTU), Singapore +3
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, leading to insufficient compositional generalization in out-of-distribution (OOD) scenarios with limited capability to capture reusable skill structures. To address this limitation, we propose Skill-Based Memory (SkillMemo) framework that implicitly decomposes long-horizon demonstrations into latent atomic skills and integrates skill-level features into a dynamic episodic memory bank for solving compositional tasks. Specifically, we first introduce an expert-guided trajectory segmentation module built upon a Mixture-of-Experts (MoE) architecture, which implicitly partitions trajectories into distinct skill primitives represented by learned gating coefficients. We further design a skill-level episodic memory architecture that stores compact skill representations as retrievable key-value pairs. During inference, the memory bank retrieves the most relevant skill primitives which are subsequently fused with the model's current gating distribution, providing a robust contextual prior to refine action predictions. Extensive experiments on the simulation benchmark and real-world manipulation tasks demonstrate that SkillMemo consistently enhances both DP and VLA backbones, achieving state-of-the-art performance and outperforming π0.5, while exhibiting strong compositional generalization to unseen task configurations.
Changyuan Wang, Chubin Zhang, Zhenyu Wu +8
Shenzhen International Graduate School, Tsinghua University · Department of Automation, Tsinghua University · Nanyang Technological University +1