Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention
Authors: Siddeshwar Raghavan, Ziqin Yuan, Fengqing Zhu, Byung-Cheol Min
Organizations: Department of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA. · Department of Computer and Information Technology,Purdue University, West Lafayette, IN, USA. · Computer Science and Intelligent Systems Engineering, Indiana University Bloomington, IN, USA.
Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills. However, retaining task performance does not ensure the behavior remains grounded in language because policies may rely on scene cues, object associations, or memorized task structure. We introduce a benchmark protocol to study how language-guided behavior changes as robotic policies learn successive tasks. We construct meaning-preserving and meaning-changing instruction variants for the Goal, Spatial, Object, and Long suites of LIBERO. Policy experiments focus on LIBERO-Goal, evaluating Original and Paraphrase instructions after each continual-learning stage. We compare representative continual imitation learning methods under their original assumptions while separating task competence from language sensitivity. The proposed diagnostics complement standard learning and forgetting metrics by measuring semantic robustness, goal adaptation, and language sensitivity. Results show that strong continual-learning performance does not always translate to reliable language grounding, and our diagnostics help determine whether retained skills remain correctly guided by their instructions. Additional materials are available at https://sites.google.com/view/stillgrounded
Figures & tables
Fig. 1: Task retention does not guarantee language grounding. Original instructions and paraphrases should lead to the same behavior, while semantic changes may require a different response. The final panel shows one possible response to an incompatible request. Our protocol evaluates these behaviors across continual learning checkpoints rather than relying only on original task success.
Fig. 2: Overview of the benchmark protocol: Original instructions and meaning-preserving paraphrases are evaluated across continual-learning stages, while one-slot minimal contrasts and scene-incompatible collisions assess final-stage goal switching and behavioral sensitivity.
Method
Original AUC ↑
Paraphrase AUC ↑
ΔAUCpara
Original Final SR ↑
Original NBT ↓
L2M
7.06 ± 0.02
5.44 ± 0.02
1.62 ± 0.04
5.01 ± 0.01
8.44 ± 0.05
SeqFT
24.63 ± 0.02
19.18 ± 0.03
5.45 ± 0.05
8.00 ± 0.02
76.49 ± 1.04
ER
58.54±0.18
51.82±0.27
6.72±0.11
59.23±0.24
17.14±0.09
EWC
32.77±0.13
27.46±0.21
5.31±0.07
27.41±0.29
46.18±0.16
PackNet
53.34±0.25
46.93±0.14
6.41±0.19
53.27±0.08
15.36±0.31
DMPEL
78.33±0.12
70.28±0.18
8.05±0.22
81.12±0.17
0.00±0.00
TABLE I: Semantic invariance on LIBERO-Goal . Metrics include Original AUC ( ↑ ), Paraphrase AUC ( ↑ ), signed paraphrase gap ΔAUCpara , Original Final SR ( ↑ ), and Original NBT ( ↓ ). We establish the standard deviation across three seeds.
Method
Target SR ↑
Average GSA ↑
Original Persistence ↓
L2M
4.42
2.81
4.61
SeqFT
7.35
4.41
8.82
ER
19.42
11.73
37.84
EWC
10.68
6.57
20.81
PackNet
15.76
8.24
42.63
DMPEL
30.91
21.68
44.72
TABLE II: Final-stage executable minimal-contrast results on LIBERO-Goal . Higher Target SR and GSA are better, while lower Original Persistence is better.
Fig. 3: Goal-Switch Accuracy (GSA) by edited semantic slot on LIBERO-Goal . Higher values indicate better adaptation to the modified instruction. No executable attribute contrasts were generated.
Fig. 4: Final-stage behavior under scene-incompatible instructions on LIBERO-Goal Left: the proportions of rollouts achieving the original goal, another audited canonical goal, or no recognized goal; vertical markers indicate the Action-Initiation Rate (AIR). Right: Language Margin versus Action Divergence, characterizing action-level sensitivity to the instruction perturbation.
Fig. 5: Distribution of the instruction perturbations used to evaluate language grounding across the four LIBERO suites. The figure summarizes the original tasks, paraphrases, generated minimal contrasts, scene-validated minimal contrasts, and scene-incompatible collision instructions, together with their semantic categories.
Fig. 6: Paraphrase sensitivity at the final stage and throughout continual learning. Final Gap measures the difference between original and paraphrase success at the final stage, whereas AUC Gap measures the corresponding difference across all stages. Lower values indicate greater semantic invariance.
Diagnostic Pair
Spearman ρ
AD vs. OGP
0.26
AIR vs. CGAR
0.80
LM vs. OGP
0.21
TABLE III: Association between action-level and outcome-level collision diagnostics across method-task pairs. Weak associations indicate that the measurements capture complementary policy behaviors.