Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-guided policy learning methods typically use an overall task description to guide the entire demonstration trajectory. For manipulation tasks involving multiple execution stages, these methods assign the same language description to different subtasks, making it difficult to distinguish the behaviors required at different stages. In this work, we propose a multi-granularity language guidance method based on instruction decomposition. The proposed method decomposes an overall task description into more fine-grained, concrete subtask-level language instructions, thereby enhancing learning efficiency and improving performance. We evaluate the proposed method in the setting of multi-task imitation learning and validate its effectiveness.
Figures & tables
Fig. 1: Comparison of language guidance strategies in imitation learning. Prior methods use a single overall instruction for the entire trajectory, while our method introduces multi-granularity language guidance by incorporating task-level instructions, fine-grained subtask-level instructions, and their combination during policy learning.
Fig. 2: Illustration of the fine-grained language annotation process. Step 1: Detect keyframes based on robot states. Step 2: Generate fine language instructions by a VLM. Step 3: Human verify to reduce annotation mess. Step 4: Assigning fine language instructions to corresponding subtask segments.
Fig. 3: Illustration of the mixed language instruction strategy. For each timestep t in a trajectory τ , a language mode m is sampled from coarse ℓc , fine ℓf , and both ℓb language instructions.
Fig. 4: Illustration of the proposed Multi-Granularity Language Guidance for Imitation Learning (MuGIL) method.
Hyperparameter
Value
Hyperparameter
Value
epochs
50
batch size
64
# encoder layers
4
# decoder layers
6
attention heads
8
action chunk size
10
history length
1
goal window sampling size
49
hidden dimension
512
image encoder
ResNet18
attention dropout
0.3
residual dropout
0.1
TABLE I: Hyperparameters of policy training in this work.
Method
Mixed Language
LSAL
LIBERO- Goal
LIBERO- Object
LIBERO- Spatial
LIBERO- 10
Avg.
MDT
63.0
88.0
73.5
48.0
68.13
MuGIL
✓
70.0
96.0
77.5
46.5
72.50
✓
✓
77.5
94.5
73.5
63.5
77.25
TABLE II: Success rates on different LIBERO task suites. The best results are shown in bold, and the second-best results are underlined.
Fig. 5: Stage-wise success counts on different LIBERO task suites.
Method
LIBERO- Goal
LIBERO- Object
LIBERO- Spatial
LIBERO- 10
Avg
MuGIL w/ Semantic Fine
77.5
94.5
73.5
63.5
77.25
MuGIL w/ Non-Semantic Fine
72.5
91.0
65.0
53.5
70.50
TABLE III: Comparison between semantic and non-semantic fine-grained language.
Method
Sampling Ratio
LIBERO-Goal
LIBERO-Object
LIBERO-Spatial
LIBERO-10
Avg.
Coarse
Fine
Both
MuGIL
0.5
0.2
0.3
77.0
94.0
68.0
53.5
73.13
0.6
0.1
0.3
77.5
94.5
73.5
63.5
77.25
0.7
0.0
0.3
74.0
86.0
69.0
56.5
71.38
TABLE IV: Different language mode sampling ratios on different LIBERO task suites.
Date pending·Jun Chen, Erdemt Bao, Wenlong Dong +7
University of Electronic Science and Technology of China · Huazhong University of Science and Technology · Southern University of Science and Technology +2