cs.ROSep 29, 2026

DROM: A Language-Guided Diffusion Framework for Multi-Skill Robotic Manipulation

Authors: Vincenzo Pomponi, Rocco Felici, Paolo Franceschi, Stefano Baraldo, Oliver Avram, Loris Roveda, Luca Maria Gambardella, Anna Valente

Organizations: Institute of Systems and Technologies for Sustainable Production (ISTePS), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland. · Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland. · Mechanical Department, Politecnico di Milano (PoliMi), Via Giuseppe Candiani, Milan, 20155, Italy. · Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Faculty of Informatics, Universit`a della Svizzera Italiana (USI), Via Buffi 13, Lugano, CH-6900, Ticino, Switzerland.

Abstract

Learning robust manipulation policies for diverse, long-horizon tasks from limited demonstrations remains a fundamental challenge in robotics. We present DROM, a language-guided diffusion framework that enables robots to learn, represent, and compose multiple manipulation skills within a single generative policy. DROM leverages Dynamic Movement Primitives (DMPs) to augment a small set of expert demonstrations into expressive multi-skill datasets, substantially reducing data collection while improving spatial generalization beyond the demonstrated workspace. Building upon Motion Planning Diffusion (MPD), we extend the diffusion architecture to support language-conditioned multi-skill trajectory generation through cross-attention, allowing a single model to generate skill-consistent motions for a diverse set of manipulation primitives, including orientation-sensitive behaviors that are difficult to design using conventional motion planning or hard-coded controllers. For long-horizon manipulation, a large language model decomposes high-level operator requests into executable sequences of skills, enabling natural language interaction and autonomous task execution. We validate DROM on a Franka Emika Panda robot, a FANUC CRX25ia robot, and in MuJoCo simulation across a wide range of manipulation tasks. Experimental results demonstrate that DROM outperforms Motion Planning Diffusion and Behavior Cloning baselines, achieves robust multi-skill generalization, and composes learned skills to reliably execute long-horizon manipulation tasks from natural language instructions using only a limited number of human demonstrations. Datasets, simulation environments, and more at https://github.com/automation-robotics-machines/drom.

Figures & tables

Explore similar work

CardsList
  1. Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation

    Dec 20, 2025Yihang Zhu, Weiqing Wang, Shijie Wu +2Robotic ManipulationStrong Imitation Learning

  2. Decompose and Reorganize: Planning with Primitives and Visuomotor Policies Learned from Demonstrations

    Jul 28, 2026Yizhou Chen, Hang Xu, Dongjie Yu +7Visuomotor PolicyImitation Learning

  3. SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

    Aug 6, 2026Changyuan Wang, Chubin Zhang, Zhenyu Wu +8Visuomotor ControlReal-World Manipulation Tasks