Purpose: Behavior cloning can accumulate errors when a learned controller visits states outside the demonstrated distribution. This study evaluates whether Disagreement-Regularized Imitation Learning (DRIL), which converts disagreement among cloned policies into a reinforcement-learning reward, improves image-based continuous control. Methods: A controlled CarRacing study combines Gaussian and Beta learner policies, demonstrations from either a clipped Gaussian expert or an intrinsically bounded Beta expert, one or 20 trajectories, deterministic and stochastic evaluation, and three retained stages: behavior cloning, the highest 10-episode training-score checkpoint, and the final DRIL checkpoint. The disagreement ensemble contains five Gaussian policies in every variant. Each retained policy is evaluated over 100 procedurally generated episodes. Results: Score-selected DRIL produced its largest gains in the few-demonstration setting, improving over the strongest behavior-cloning mean by 61% with clipped-action demonstrations and by 112% with bounded-action demonstrations. With 20 trajectories, the advantage of DRIL narrowed; in the bounded-action regime, Beta behavior cloning remained about 7% above the best DRIL checkpoint. The experiments also show that the informativeness of the disagreement reward changes with the learner representation and training stage. Conclusion: DRIL can substantially improve few-demonstration visual continuous control, while bounded Beta policies provide strong behavior-cloning performance when more demonstrations are available. The results highlight the joint importance of learner support,ensemble response, and checkpoint selection.
Figures & tables
Method
Expert queries
Task reward
Corrective mechanism
BC
No
No
None; supervised fitting on the fixed dataset
DAgger Ross et al. (2011)
Yes
No
Aggregates expert labels on learner-visited states
GAIL Ho and Ermon (2016)
No
No
Adversarial occupancy-measure matching
Hybrid RL–IL Kumar (2024) ; Lu et al. (2023)
Varies
Yes
Refines imitation with environment rewards
Diffusion Meets DAgger Zhang et al. (2024)
No new labels
No
Synthesizes examples near failure states
DRIL Brantley et al. (2020)
No
No
Thresholded disagreement among cloned policies
Table 1 : Qualitative comparison of methods that address imitation-learning distribution shift. “Expert queries” indicates whether additional labels are required after the initial demonstrations. “Task reward” indicates use of the environment reward during policy improvement.
Figure 1 : Two-stage DRIL workflow used in this work. In stage (a), the expert demonstration dataset D trains the ensemble Π by behavior cloning; in the present experiments, Π contains five Gaussian policies trained on bootstrap samples. In stage (b), a BC-initialized learner interacts with the environment. Ensemble disagreement provides the binary reward for reinforcement-learning updates, while behavior-cloning updates continue to use the original demonstrations.
Figure 2 : Visual actor–critic architectures used in this work. Both learners use the same image preprocessing and convolutional feature extractor. The actor head changes from Gaussian outputs in (a) to Beta-distribution parameters in (b), while the critic estimates a scalar value function. The disagreement ensemble remains Gaussian in all experiments.
Figure 3 : Expert-action generation. The Beta expert produces bounded samples; the Gaussian expert produces many samples outside the action range; the clipped Gaussian actions are those executed in the environment and used in the clipped-action demonstration dataset.
Factor
Levels
Expert demonstration actions
Clipped Gaussian; intrinsically bounded Beta
Number of expert trajectories
1; 20
Learner policy distribution
Gaussian; Beta
Retained training stage
BC; DRIL-peak; DRIL-final
Evaluation mode
Deterministic mean; stochastic sampling
Table 2 : Experimental factors in the main CarRacing study. Every learner is evaluated in deterministic mode, using the distribution mean, and stochastic mode, sampling from the distribution. BC denotes the validation-selected behavior-cloning checkpoint; DRIL-peak is selected by the highest 10-episode training-score mean; DRIL-final is the last checkpoint.
Parameter
Value
Parameter
Value
BC learning rate
2.5×10−4
BC mini-batch
32
BC data split
80/20
BC patience
20 epochs
Ensemble size
5
Ensemble training
2000 epochs
Uncertainty quantile
0.98
PPO learning rate
3×10−4 , annealed
PPO rollout
2048 steps
PPO epochs
10
PPO mini-batches
32
Discount γ
0.99
Table 3 : Main CarRacing training settings. Ensemble members are Gaussian policies trained for 2000 epochs on bootstrap samples. The learner is first trained by BC and then updated by interleaved PPO and BC steps during DRIL.
Expert
Method
1 trajectory
20 trajectories
Deterministic
Stochastic
Deterministic
Stochastic
Clipped
BC
125 ± 113
30 ± 67
171 ± 124
437 ± 115
DRIL-peak
166 ± 131
322 ± 208
423 ± 211
802 ± 197
DRIL-final
41 ± 80
39 ± 79
229 ± 107
218 ± 108
Bounded
BC
194 ± 113
75 ± 47
617 ± 260
137 ± 70
DRIL-peak
235 ± 123
341 ± 246
656 ± 266
720 ± 289
Table 4 : Gaussian learner evaluation on CarRacing. Values are mean ± standard deviation over 100 evaluation episodes. Bold values mark the highest mean scores within each evaluation mode and dataset regime.
Figure 4 : Training dynamics of the Gaussian learner using clipped-action expert demonstrations. Panels (a) and (b) correspond to one and 20 demonstrations, respectively.
Figure 5 : Stochastic 100-episode evaluation of the Gaussian learner using clipped-action expert demonstrations. Panels (a) and (b) correspond to one and 20 demonstrations, respectively, and compare BC, DRIL-peak, and DRIL-final. Deterministic results are reported in Table 4 .
Figure 6 : Training dynamics of the Gaussian learner using bounded-action expert demonstrations. Panels (a) and (b) correspond to one and 20 demonstrations, respectively.
Figure 7 : Stochastic 100-episode evaluation of the Gaussian learner using bounded-action expert demonstrations. Panels (a) and (b) correspond to one and 20 demonstrations, respectively, and compare BC, DRIL-peak, and DRIL-final. Deterministic results are reported in Table 4 .
Expert
Method
1 trajectory
20 trajectories
Deterministic
Stochastic
Deterministic
Stochastic
Clipped
BC
163 ± 67
216 ± 108
252 ± 143
758 ± 237
DRIL-peak
229 ± 142
348 ± 230
210 ± 115
572 ± 280
DRIL-final
122 ± 116
143 ± 111
265 ± 165
241 ± 140
Bounded
BC
147 ± 105
161 ± 156
567 ± 239
794 ± 227
DRIL-peak
195 ± 88
200 ± 184
398 ± 225
738 ± 277
Table 5 : Beta learner evaluation on CarRacing with the fixed Gaussian disagreement ensemble. Values are mean ± standard deviation over 100 evaluation episodes. Bold values mark the highest mean scores within each evaluation mode and dataset regime.
Figure 8 : Training dynamics of the Beta learner using clipped-action expert demonstrations and the fixed Gaussian disagreement ensemble. Panels (a) and (b) correspond to one and 20 demonstrations, respectively.
Figure 9 : Stochastic 100-episode evaluation of the Beta learner using clipped-action expert demonstrations and the fixed Gaussian disagreement ensemble. Panels (a) and (b) correspond to one and 20 demonstrations, respectively, and compare BC, DRIL-peak, and DRIL-final. Deterministic results are reported in Table 5 .
Figure 10 : Training dynamics of the Beta learner using bounded-action expert demonstrations and the fixed Gaussian disagreement ensemble. Panels (a) and (b) correspond to one and 20 demonstrations, respectively.
Figure 11 : Stochastic 100-episode evaluation of the Beta learner using bounded-action expert demonstrations and the fixed Gaussian disagreement ensemble. Panels (a) and (b) correspond to one and 20 demonstrations, respectively, and compare BC, DRIL-peak, and DRIL-final. Deterministic results are reported in Table 5 .
Expert
Trajectories
Best BC
Best DRIL-peak
Change of DRIL peak
Clipped
1
Beta: 216 ± 108
Beta: 348 ± 230
+61.1%
Clipped
20
Beta: 758 ± 237
Gaussian: 802 ± 197
+5.8%
Bounded
1
Beta: 161 ± 156
Gaussian: 341 ± 246
+111.8%
Bounded
20
Beta: 794 ± 227
Beta: 738 ± 277
−7.1 %
Table 6 : Best stochastic BC and score-selected DRIL checkpoint in each demonstration regime, allowing the learner distribution to differ. Percent change is computed from the reported means.
Figure 12 : Disagreement signal during the first 1000 interaction steps of the Gaussian learner. The left column uses one expert trajectory and the right column uses 20 trajectories. The top rows show steering and brake/throttle samples with the valid action interval shaded; the middle rows show scalar ensemble uncertainty and the 98th-percentile threshold; the bottom rows show the resulting binary reward. Sampled Gaussian actions frequently extend beyond the valid interval before clipping, and occasional uncertainty values cross the threshold.
Figure 13 : Disagreement signal during the first 1000 interaction steps of the Beta learner using the same fixed Gaussian ensemble. The left column uses one expert trajectory and the right column uses 20 trajectories. The top rows show bounded steering and brake/throttle actions; the middle rows show scalar ensemble uncertainty and the 98th-percentile threshold; the bottom rows show the binary reward. Threshold crossings are infrequent, so the reward remains positive over most of the rollout.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Environment
Demonstrations
Expert threshold
BC
DRIL
Breakout
1 trajectory
300
5
355
Breakout
3 trajectories
300
7
339
LunarLanderContinuous
1 trajectory
200
194
257
LunarLanderContinuous
20 trajectories
200
232
228
Appendix
Table 7 : Implementation checks in the two regimes considered by the original DRIL study. Each entry is the mean score over 100 evaluation episodes. The expert threshold is the task score used to define successful or expert-level performance in the corresponding experiment.
Figure 14 : Implementation-validation results. The figures show the qualitative behavior of DRIL in the original discrete-visual and continuous-low-dimensional regimes before the image-based continuous-control study.
Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment. We show that reusing the training data during inference via a semi-parametric retrieval-based imitation learning approach can alleviate this challenge. We present Difference-Aware Retrieval Policies for Imitation Learning (DARP), a semi-parametric retrieval-based imitation learning approach that addresses this limitation by reparameterizing the imitation learning problem in terms of local neighborhood structure rather than direct state-to-action mappings. Instead of learning a global policy, DARP trains a model to predict actions based on k-nearest neighbors from expert demonstrations, their corresponding actions, and the relative distance vectors between neighbor states and query states. DARP requires no additional assumptions beyond those made for standard behavior cloning -- it does not require additional data collection, online expert feedback, or task-specific knowledge. We demonstrate consistent performance improvements of 15-46% over standard behavior cloning across diverse domains, including continuous control and robotic manipulation, and across different representations, including high-dimensional visual features. Code and demos are available at https://weirdlabuw.github.io/darp-site/.
Quinn Pfeifer, Ethan Pronovost, Paarth Shah +3
1Paul G. Allen School of Computer Science & Engineering, University of Washington · 2Toyota Research Institute · 3Google DeepMind +1
Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions. By blending human and policy actions, we aim to improve the autonomous performance of manipulation policies. We validate our approach across five manipulation tasks, two in the real world and three in simulation. Our approach achieves higher autonomous performance by 30 or more percentage points on two real-world tasks compared to a typical human-gated correction approach (HG-DAgger). We also investigate the advantages of BlenDAgger that allow for higher autonomous performance, finding that BlenDAgger results in 57% smoother transitions between policy control and human interventions, and 14% higher trajectory similarity to the training data. In a user study (n=14) on two real-world tasks, we find that BlenDAgger results in faster data collection (BF=13.32), and we do not find a difference in subjective perceptions. These results show that blended shared control leads to higher autonomous performance compared to typical methods for fine-tuning robot policies from fully teleoperated interventions.
Cailyn Smith, Geoffrey Sun, Henny Admoni +1
The Robotics Institute, School of Computer Science Carnegie Mellon University Pittsburgh, USA · School of Computer Science Carnegie Mellon University Pittsburgh, USA
Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights. These scalar signals are inherently limited, as they cannot explicitly express intermediate reasoning about task progress, failure modes, or corrective actions. We propose a language-critique framework for imitation learning from suboptimal demonstrations that instead leverages natural language as a structured supervision signal, avoiding the collapse of expressive feedback into scalars. Our method first constructs language labels from demonstrations that explicitly describe current progress, identify suboptimal behaviors, and provide fine-grained corrective guidance. We then introduce a language-critique loss that directly trains policies using these structured signals without reducing them to scalars, and instantiate it for both behavior cloning and diffusion policies, yielding LC-BC and LC-DP. We further provide a theoretical result showing that the proposed objective upper-bounds the expert performance gap under standard assumptions. Empirically, we evaluate on diverse continuous control tasks spanning navigation, manipulation, and gameplay, where our methods consistently outperform strong imitation learning and offline reinforcement learning baselines. These results demonstrate that language can serve as a powerful and structured form of supervision for learning robust policies from suboptimal data.
Chih-Han Yang, Dai-Jie Wu, Yun-Ping Huang +3
Graduate Institute of Communication Engineering, National Taiwan University (NTU) · University of Utah · National Yang Ming Chiao Tung University +1