DROM: A Language-Guided Diffusion Framework for Multi-Skill Robotic Manipulation
Authors: Vincenzo Pomponi, Rocco Felici, Paolo Franceschi, Stefano Baraldo, Oliver Avram, Loris Roveda, Luca Maria Gambardella, Anna Valente
Organizations: Institute of Systems and Technologies for Sustainable Production (ISTePS), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland. · Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland. · Mechanical Department, Politecnico di Milano (PoliMi), Via Giuseppe Candiani, Milan, 20155, Italy. · Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Faculty of Informatics, Universit`a della Svizzera Italiana (USI), Via Buffi 13, Lugano, CH-6900, Ticino, Switzerland.
Learning robust manipulation policies for diverse, long-horizon tasks from limited demonstrations remains a fundamental challenge in robotics. We present DROM, a language-guided diffusion framework that enables robots to learn, represent, and compose multiple manipulation skills within a single generative policy. DROM leverages Dynamic Movement Primitives (DMPs) to augment a small set of expert demonstrations into expressive multi-skill datasets, substantially reducing data collection while improving spatial generalization beyond the demonstrated workspace. Building upon Motion Planning Diffusion (MPD), we extend the diffusion architecture to support language-conditioned multi-skill trajectory generation through cross-attention, allowing a single model to generate skill-consistent motions for a diverse set of manipulation primitives, including orientation-sensitive behaviors that are difficult to design using conventional motion planning or hard-coded controllers. For long-horizon manipulation, a large language model decomposes high-level operator requests into executable sequences of skills, enabling natural language interaction and autonomous task execution. We validate DROM on a Franka Emika Panda robot, a FANUC CRX25ia robot, and in MuJoCo simulation across a wide range of manipulation tasks. Experimental results demonstrate that DROM outperforms Motion Planning Diffusion and Behavior Cloning baselines, achieves robust multi-skill generalization, and composes learned skills to reliably execute long-horizon manipulation tasks from natural language instructions using only a limited number of human demonstrations. Datasets, simulation environments, and more at https://github.com/automation-robotics-machines/drom.
Figures & tables
Figure 1 : Illustration of the two operational modalities of DROM: (a) language-guided single-skill execution, where an operator query is embedded and mapped to a manipulation primitive to generate a skill-consistent trajectory; and (b) language-guided long-horizon planning, where a language model decomposes a high-level objective into an ordered sequence of skills, each executed via diffusion-based trajectory generation.
Figure 2 : Overview of the diffusion-based training pipeline of DROM: (A) collection of a single expert demonstration per skill; (B) synthetic dataset expansion using Dynamic Movement Primitives (DMPs); (C) training of the skill-conditioned diffusion model on the augmented multi-skill dataset; and (D) deployment in simulation and on real robotic platforms.
Figure 3 : The set of manipulation skills considered in this study for the FANUC CRX25ia robot (a–i) and the Franka Emika Panda (j–o).
Skill
Demos
Skill
Demos
Co-Transport
1
Pick Battery
2
Shelf
1
Place Battery
2
Slide
1
Pick Bottle
4
Stir
1
Place Bottle
4
Pick Cube
1
Pick Cup
4
Stack Cube
1
Pour Liquid
4
Table 1 : Number of demonstrations collected for each manipulation skill.
Figure 4 : Architecture of the proposed skill-conditioned diffusion model. At each reverse diffusion step, a Temporal U-Net denoiser predicts the noise of the trajectory conditioned on the diffusion timestep, the environment state, and the language embedding of the user prompt. Cross-attention layers at every level of the U-Net allow the model to attend to skill-specific semantic information while preserving temporal structure through the U-Net backbone.
Skill
1 Demo
2 Demos
4 Demos
Pick Battery
48%
92%
92%
Place Battery
50%
98%
98%
Pick Bottle
24%
72%
80%
Place Bottle
24%
78%
84%
Pick Cup
28%
46%
86%
Pour Liquid
22%
56%
92%
Table 2 : Effect of the number of demonstrations on Diffusion Model performance. Success rates obtained for orientation-sensitive manipulation skills when increasing the number of demonstrations provided to the DMP augmentation process.
Object
Px
Py
Pz
Rz
Battery
23%
21%
7%
16%
Bottle
24%
24%
8%
9%
Button
24%
24%
8%
5%
Co-Transport Object
28%
22%
7%
23%
Pouring Container
31%
19%
3%
19%
Shelf
17%
29%
2%
33%
Table 3 : Workspace of each manipulated object, reported in terms of translational and rotational variations.
FANUC CRX25ia
Franka Research 3
Skill
MPD
DROM
Skill
MPD
DROM
Pick Battery
76%
98%
Pick Cube
78%
90%
Place Battery
82%
96%
Stack Cube
44%
94%
Pick Bottle
60%
84%
Grab Handle
38%
90%
Place Bottle
76%
94%
Open Drawer
50%
86%
Pick Cup
38%
86%
Place in Drawer
62%
92%
Table 4 : Comparison between DROM and MPD on the considered manipulation skills. The left half reports results on the FANUC CRX25ia, while the right half reports results on the Franka Emika Panda. DROM consistently outperforms MPD across both robotic platforms by learning multiple manipulation skills within a single language-conditioned diffusion model.
Figure 5 : Simulation environments. The two environments are designed to closely replicate the manipulation setups executed on the real Franka arm, while introducing structured, multi-stage objectives that decompose naturally into sequential subtasks.
Env.
Skill
MPD
BC
DROM
MugCleanup
Pick Cube
59.33±2.49
96.00±2.00
97.33±1.15
Stack Cube
60.67±0.94
92.67±2.31
94.67±0.94
Stack
Grab Handle
28.00±1.63
92.00±3.46
96.00±1.63
Open Drawer
58.00±1.63
90.67±0.94
96.67±0.94
Pick Mug
28.00±2.83
85.33±4.11
88.00±1.63
Place in Drawer
30.67±0.94
92.00±1.63
94.67±2.49
Table 5 : Comparison of DROM, MPD, and BC, in the MuJoCo simulation environments: Stack and Mug Cleanup.
Figure 6 : Qualitative examples of multimodal trajectory generation. (a) Pick Bottle : depending on the bottle orientation, the diffusion model generates different approach trajectories to insert a gripper finger into the handle. (b) Grab Handle : the model adapts the grasping motion to two drawers with different handle geometries, producing the corresponding manipulation strategy.
Task
# Skills
# Objects
SR
Pick and Place
2
2
86.0%
Pour
2
1
70.0%
Stack
2
2
67.5%
Hide
5
1
70.4%
Clean Up
5 – 9
1 – 3
60.9%
Table 6 : End-to-end evaluation of DROM. Success rate (SR) achieved on long-horizon manipulation tasks of increasing complexity.
Task
Skill Composition
Sequence Execution
Overall
Pick and Place
100%
86.0%
86.0%
Pour
87.5%
80.0%
70.0%
Stack
75.0%
90.0%
67.5%
Hide
80.0%
88.0%
70.4%
Clean Up
72.5%
84.0%
60.9%
Table 7 : Success rates of skill composition, sequence execution and overall for Stack and Clean Up tasks.
Category
Failure Mode
#
%
Language Model
Incorrect skill sequence
42/250
16.8
Diffusion Model
Failed grasp
11/250
4.4
Placement failure
7/250
2.8
Collision
7/250
2.8
Unfeasible trajectory
1/250
0.4
Goal not reached
10/250
4
Table 8 : Failure analysis of DROM. Failures are grouped according to the corresponding module of the framework.
Skill
Precision
Recall
F1-score
SR
Pick Cube
96.5%
83.0%
89.2%
89.0%
Stack Cube
97.6%
82.0%
89.1%
85.0%
Grab Handle
74.6%
85.0%
79.4%
82.0%
Open Drawer
85.6%
89.0%
87.3%
85.0%
Place in Drawer
75.0%
93.0%
83.0%
93.0%
Close Drawer
96.6%
85.0%
90.4%
83.0%
Table 9 : Language encoder evaluation on a balanced test set of 600 natural language instructions. Precision, recall, F1-score, and class-wise success rate (SR) quantify the encoder’s ability to associate operator requests with the corresponding manipulation primitive.
Figure 7 : Confusion matrix of the language encoder on the six manipulation primitives executed with the Franka Emika Panda robot.
Skill
Language Recognition
Imitation Learning
Overall Success
Pick Cube
89.0%
90.0%
80.1%
Stack Cube
85.0%
94.0%
79.9%
Grab Handle
82.0%
90.0%
73.8%
Open Drawer
85.0%
86.0%
73.1%
Place in Drawer
93.0%
92.0%
85.6%
Close Drawer
83.0%
88.0%
73.0%
Table 10 : Evaluation of language-guided diffusion planning , in terms of language recognition accuracy, skill execution success rate, and request-to-goal success rate.