Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-guided policy learning methods typically use an overall task description to guide the entire demonstration trajectory. For manipulation tasks involving multiple execution stages, these methods assign the same language description to different subtasks, making it difficult to distinguish the behaviors required at different stages. In this work, we propose a multi-granularity language guidance method based on instruction decomposition. The proposed method decomposes an overall task description into more fine-grained, concrete subtask-level language instructions, thereby enhancing learning efficiency and improving performance. We evaluate the proposed method in the setting of multi-task imitation learning and validate its effectiveness.
Figures & tables
Fig. 1: Comparison of language guidance strategies in imitation learning. Prior methods use a single overall instruction for the entire trajectory, while our method introduces multi-granularity language guidance by incorporating task-level instructions, fine-grained subtask-level instructions, and their combination during policy learning.
Fig. 2: Illustration of the fine-grained language annotation process. Step 1: Detect keyframes based on robot states. Step 2: Generate fine language instructions by a VLM. Step 3: Human verify to reduce annotation mess. Step 4: Assigning fine language instructions to corresponding subtask segments.
Fig. 3: Illustration of the mixed language instruction strategy. For each timestep t in a trajectory τ , a language mode m is sampled from coarse ℓc , fine ℓf , and both ℓb language instructions.
Fig. 4: Illustration of the proposed Multi-Granularity Language Guidance for Imitation Learning (MuGIL) method.
Hyperparameter
Value
Hyperparameter
Value
epochs
50
batch size
64
# encoder layers
4
# decoder layers
6
attention heads
8
action chunk size
10
history length
1
goal window sampling size
49
hidden dimension
512
image encoder
ResNet18
attention dropout
0.3
residual dropout
0.1
TABLE I: Hyperparameters of policy training in this work.
Method
Mixed Language
LSAL
LIBERO- Goal
LIBERO- Object
LIBERO- Spatial
LIBERO- 10
Avg.
MDT
63.0
88.0
73.5
48.0
68.13
MuGIL
✓
70.0
96.0
77.5
46.5
72.50
✓
✓
77.5
94.5
73.5
63.5
77.25
TABLE II: Success rates on different LIBERO task suites. The best results are shown in bold, and the second-best results are underlined.
Fig. 5: Stage-wise success counts on different LIBERO task suites.
Method
LIBERO- Goal
LIBERO- Object
LIBERO- Spatial
LIBERO- 10
Avg
MuGIL w/ Semantic Fine
77.5
94.5
73.5
63.5
77.25
MuGIL w/ Non-Semantic Fine
72.5
91.0
65.0
53.5
70.50
TABLE III: Comparison between semantic and non-semantic fine-grained language.
Method
Sampling Ratio
LIBERO-Goal
LIBERO-Object
LIBERO-Spatial
LIBERO-10
Avg.
Coarse
Fine
Both
MuGIL
0.5
0.2
0.3
77.0
94.0
68.0
53.5
73.13
0.6
0.1
0.3
77.5
94.5
73.5
63.5
77.25
0.7
0.0
0.3
74.0
86.0
69.0
56.5
71.38
TABLE IV: Different language mode sampling ratios on different LIBERO task suites.
Real-world robotic disassembly requires long-horizon execution, where robots must perform ordered sequences of manipulation tasks across multiple parts within a single scene. Multiple valid task goals and diverse assembly configurations make it difficult for imitation policies to infer the intended skill from raw observations alone, particularly when training data cannot cover the combinatorial diversity of real-world configurations and part geometries. We show that incorporating task context through language alleviates these challenges by providing explicit structure for skill selection and associating language-specified tasks with their corresponding manipulation targets in the visual scene. The proposed framework combines hierarchical task selection with task-context-aware imitation learning to ground language instructions in spatial visual representations for robotic disassembly. The resulting framework generalizes across diverse connector geometries and assembly configurations without requiring explicit object annotations. Our method improves end-to-end task success by 35 percentage points over the baseline diffusion policy and by 75 percentage points over the previous task-context-aware baseline.
Jeon Ho Kang, Igal Tamarkin, Ethan Niu +2
Viterbi School of Engineering, University of Southern California, Los Angeles, USA
Language-conditioned Imitation Learning (IL) is essential for enabling robots to perform complex tasks following natural language instructions. However, generalizing to multi-step compositional tasks remains a significant challenge. While hierarchical approaches attempt to address this by decomposing tasks into atomic skills, existing methods often suffer from training instability and codebook collapse due to the tight coupling between high-level skill reasoning and low-level action generation in joint training paradigms. Inspired by the Dual-Process Theory of cognition, we propose Dual-Process Atomic Skill Learning (DASL), a novel asynchronous hierarchical imitation learning framework that decouples slow semantic reasoning from fast, real-time motion control. DASL comprises a Slow-Frequency Policy that predicts interpretable, discrete skills via Vector Quantization, and a High-Frequency Policy that leverages a latent diffusion model and a Decision Transformer to generate precise actions conditioned on these latent skills. By asynchronously coordinating these modules and utilizing diffusion to structure the latent space, our framework mitigates the skill codebook interference problem common in joint training paradigms. Evaluations across simulation benchmarks and experiment demonstrate that DASL significantly outperforms state-of-the-art baselines, excelling in skill acquisition and compositional generalization to unseen instructions. GitHub page: https://github.com/Hatakekaka/DASL
Jun Chen, Erdemt Bao, Wenlong Dong +7
University of Electronic Science and Technology of China · Huazhong University of Science and Technology · Southern University of Science and Technology +2
Instruction granularity is an important yet poorly controlled variable in language-guided embodied AI. Existing benchmarks typically pair each task with a single static instruction, making it difficult to study how agent behavior changes when the same task is described at different levels of detail. We introduce Mini-BEHAVIOR-Gran, a new benchmark for controlled studies of instruction granularity that extends Mini-BEHAVIOR with multiple instruction variants per task, ranging from high-level goal descriptions to step-by-step guidance. Using this benchmark, we compare four candidate metrics for cross-task granularity quantification: token count, entity count, action-verb count, and planning-width, and find that width correlates most consistently with agent performance. Using width to organize training and evaluation further reveals a non-monotonic U-shaped relationship between instruction granularity and performance, with peaks at both fine and coarse extremes. Further analysis suggests that the coarse-granularity performance rebound is associated with shallow grounding, where agents learn vision-dominant policies.
Sukai Huang, Chenyuan Zhang, Fucai Ke +4
Faculty of Information Technology, Monash University