cs.LGSep 29, 2026

Disagreement-Regularized Imitation Learning for Image-Based Continuous Control with Gaussian and Beta Policies

Authors: Irving Giovani Bronzatti Petrazzini, Eric Aislan Antonelo

Organizations: Department of Automation and Systems Engineering, Federal University of Santa Catarina, Florian´opolis, Santa Catarina, Brazil.

Abstract

Purpose: Behavior cloning can accumulate errors when a learned controller visits states outside the demonstrated distribution. This study evaluates whether Disagreement-Regularized Imitation Learning (DRIL), which converts disagreement among cloned policies into a reinforcement-learning reward, improves image-based continuous control. Methods: A controlled CarRacing study combines Gaussian and Beta learner policies, demonstrations from either a clipped Gaussian expert or an intrinsically bounded Beta expert, one or 20 trajectories, deterministic and stochastic evaluation, and three retained stages: behavior cloning, the highest 10-episode training-score checkpoint, and the final DRIL checkpoint. The disagreement ensemble contains five Gaussian policies in every variant. Each retained policy is evaluated over 100 procedurally generated episodes. Results: Score-selected DRIL produced its largest gains in the few-demonstration setting, improving over the strongest behavior-cloning mean by 61% with clipped-action demonstrations and by 112% with bounded-action demonstrations. With 20 trajectories, the advantage of DRIL narrowed; in the bounded-action regime, Beta behavior cloning remained about 7% above the best DRIL checkpoint. The experiments also show that the informativeness of the disagreement reward changes with the learner representation and training stage. Conclusion: DRIL can substantially improve few-demonstration visual continuous control, while bounded Beta policies provide strong behavior-cloning performance when more demonstrations are available. The results highlight the joint importance of learner support,ensemble response, and checkpoint selection.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Difference-Aware Retrieval Policies for Imitation Learning

    Jun 8, 2026Quinn Pfeifer, Ethan Pronovost, Paarth Shah +3Imitation LearningBehavior Cloning

  2. BlenDAgger: Blended Shared Control for Interactive Imitation Learning

    Sep 29, 2026Cailyn Smith, Geoffrey Sun, Henny Admoni +1Robotic Manipulation PoliciesRobot Policies

  3. Language-Critique Imitation Learning from Suboptimal Demonstrations

    Jul 1, 2026Chih-Han Yang, Dai-Jie Wu, Yun-Ping Huang +3Temporally Coherent Imitation LearningStrong Imitation Learning