Organizations: The Robotics Institute, School of Computer Science Carnegie Mellon University Pittsburgh, USA · School of Computer Science Carnegie Mellon University Pittsburgh, USA
Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions. By blending human and policy actions, we aim to improve the autonomous performance of manipulation policies. We validate our approach across five manipulation tasks, two in the real world and three in simulation. Our approach achieves higher autonomous performance by 30 or more percentage points on two real-world tasks compared to a typical human-gated correction approach (HG-DAgger). We also investigate the advantages of BlenDAgger that allow for higher autonomous performance, finding that BlenDAgger results in 57% smoother transitions between policy control and human interventions, and 14% higher trajectory similarity to the training data. In a user study (n=14) on two real-world tasks, we find that BlenDAgger results in faster data collection (BF=13.32), and we do not find a difference in subjective perceptions. These results show that blended shared control leads to higher autonomous performance compared to typical methods for fine-tuning robot policies from fully teleoperated interventions.
Figures & tables
Fig. 1 : BlenDAgger shares control between the robot’s policy and the human demonstrator during corrections, resulting in higher autonomous success and faster data collection.
Fig. 2 : BlenDAgger pipeline. When the demonstrator chooses to intervene, their action is blended with the policy’s predicted action. The resulting blended action is executed on the robot and used to fine-tune the policy.
Fig. 3 : Experiment Tasks.
Fig. 4 : α over an example 30-second segment. Highlighted sections show timesteps where the participant is intervening.
Fig. 5 : Autonomous performance on real-world (a–b) and simulation (c–e) tasks. Bar colors indicate the highest subtask reached during an episode. Full task success is labeled.
Simulation
Real
Architecture
External / wrist cameras
2 / 1
1 / 1
Image resolution
128×128
Random crop
116×116
112×112
Proprioception dim.
16
14
Action dim.
12
10
TABLE I : Policy architecture and training hyperparameters.
Object
Robot base
RoboCasa
Task
Object
x / y Pos. (cm)
Yaw (rad)
x / y Pos. (cm)
Yaw (rad)
Layout
Style
Prepare Coffee
Mug
± 10 / 10
0
± 2.5 / 4
±0.17
7
9
Microwave Thawing
Carrot
± 5 / 5
0
± 1.5 / 4
±0.17
4
0
Pick and Place
Lemon
± 10 / 10
2π
0
0
1
1
Cupboard Stowing
Can
± 11 / 11
2π
0
0
—
—
Cupboard
0 / 0
π/6
0
0
—
—
TABLE II : Randomization for data collection and evaluation.
Fig. 6 : Participants complete the task faster and with fewer intervention steps with BlenDAgger. Error bars show 95% CIs.
Fig. 7 : Autonomous performance (user study data).
Fig. 8 : BlenDAgger has smoother action transitions and lower trajectory distance to training data. Error bars show 95% CIs.
Metric
Experts
Participants
Action discontinuity, intervention start/end
6.2e7 / 1.5e6
15.34 / 3.8e36
Highly OOD states visited
0.34
17.07
Trajectory distance to training data
3.35
3.10
Intervention timing, intervention start/end
0.51 / 0.53
0.99 / 1.66
TABLE III : BFs comparing BlenDAgger to HG-DAgger on analysis metrics. Bold indicates evidence of an effect.
Performance
Responsive-
Method
R1
R2
R3
ness (s)
HG-DAgger
0.01 ± 0.00
0.02 ± 0.01
0.05 ± 0.01
0.15
BlenDAgger
0.08 ± 0.03
0.15 ± 0.02
0.27 ± 0.04
0.20
Without uncertainty
0.00
0.14
0.16
0.28
Without similarity
0.01
0.15
0.22
0.20
Without magnitude
0.03
0.13
0.17
0.20
TABLE IV : Ablation performance and responsiveness. Bold marks the best overall and the best BlenDAgger variant.
Learning from demonstrations is effective for robotic manipulation, but collecting sufficient task-specific data remains a major bottleneck. Under distribution shift, small errors compound, performance degrades, and expert time is often spent on redundant, low-value corrections instead of the few critical failure cases. We present VR-DAgger, a human-in-the-loop framework centered on an immersive VR application for dexterous teleoperation, demonstration collection, and selective policy correction. The VR client provides intuitive hand control with synchronized scene visualization, while a backend workstation runs simulation and learning, enabling autonomous rollouts without continuous operator oversight. We use Monte Carlo (MC) dropout to score uncertainty during Isaac Lab rollouts of a diffusion policy and select informative failure segments for correction. These segments are replayed in VR as clips, where the operator selectively labels and corrects the policy's behavior, concentrating supervision where uncertainty is highest without full-rollout monitoring or a separate intervention classifier. We evaluate on three dexterous manipulation tasks (Pan pick-and-place, Drawer opening, Valve turning) with a 10-DoF XHand under standard and challenging initial configurations. Active labeling consistently improves over behavioral cloning across all tasks, with gains of up to 23 percentage points. Compared to unguided human-in-the-loop inspection, VR-DAgger reduces per-sample collection time by approximately 40% by focusing review on selected segments rather than full rollouts.
René Zurbrügg, Tifanny Portela, Arjun Bhardwaj +3
The Authors are with the Robotics Systems Lab ETH Zürich.
Large behaviour models have transformed the field of robotic manipulation, but prohibitive data requirements have thus far prevented a revolution similar to vision language models. We believe that instrumentation, i.e. sensor integration in objects, can provide invaluable state information and enable efficient learning for robotic manipulation. In this paper, we present instrumented imitation learning of clothes hanger insertion. Using 180 teleoperated demonstrations, we train diffusion policies with and without access to instrumentation data. Results show that policies leveraging instrumentation outperform vision-only counterparts by 14-25 %pt and exhibit greater task awareness. Crucially, a black-box imitation learning policy learns to prioritise instrumentation signals without explicit guidance. In addition, enhancing the teleoperation dataset with rollouts from an instrumented expert policy, enables a vision-only student policy to achieve performance comparable to the instrumented expert, thereby surpassing the original vision-only policy. These findings establish instrumentation as a promising strategy to enhance imitation learning for robotic manipulation. Datasets are available on Zenodo.
Remko Proesmans, Thomas Lips, Francis wyffels
AI and Robotics Lab (IDLab-AIRO), Ghent University—imec, Ghent, Belgium
Behavior cloning for robot manipulation relies on expert demonstrations. However, for tasks that require dynamic stability, precise contact timing, or dexterous coordination, human operators may find it hard or even impossible to collect data. We study this infeasible-demonstration regime and propose GLIDE: Guardrails for Learning from Infeasible Demonstrations Efficiently, a framework that infers task-specific failure modes and converts them into executable guardrails for data collection and policy deployment. Given a task description and the conditioning teleoperation code, GLIDE writes guardrails that use system states to filter teleoperation and policy commands, constrain failure-prone actions, and iteratively improve from trajectory feedback. Across three tasks, GLIDE discovers emergent guardrails that go beyond domain-expert hardcoded ones, improving data collection over naive VR teleoperation and domain-expert hardcoded guardrails. After refinement, GLIDE raises data-collection success from 0-10 percent to 70-90 percent across the three tasks. During policy execution, mixed-data guarded policies reach 70 percent, 60 percent, and 60 percent success on Tomato plate transfer, Marker handover and stand, and Wine serving tasks. These results show that GLIDE can support policy learning when direct demonstrations are infeasible. Project website: http://guardrail-policy.github.io/