Organizations: Division of Information Science, Graduate School of Information Science, Nara Institute of Science and Technology (NAIST), Nara, Japan · Department of Industrial Engineering of the University of Trento, Trento, Italy
With the increasing use of robot-free demonstration interfaces that provide state trajectories without action labels, imitation from observation has become a promising approach for learning robot behaviors from human demonstrations. However, due to differences in embodiment and dynamics between humans and robots, demonstrated human motions may not be feasible for the robot, potentially degrading policy performance. In this study, we propose Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO), which estimates the feasibility of state-only demonstrations from the robot's own experience rather than relying on explicit dynamics models or large prior exploration datasets. A key feature of EF-GAIfO is that the notion of feasibility evolves with policy learning: as the policy improves and the robot experiences a broader range of state transitions, the feasible region is progressively expanded, allowing additional demonstrations to be incorporated into learning. This enables feasibility-aware imitation that adapts to the current stage of policy learning, rather than relying on a pre-designed feasibility criterion. We validate the effectiveness of EF-GAIfO on a locomotion task in simulation and on a real quadruped robot performing a object-reaching-and-grasping task.
Figures & tables
Fig. 1 : Evolution of feasibility estimation during policy learning. Early in learning, limited robot experience yields high feasibility estimates only for demonstration transitions near the experienced region. As experience accumulates, the covered region expands, allowing more demonstration transitions to receive high feasibility estimates. This progressively reduces the influence of infeasible demonstrations during imitation learning.
Fig. 2 : Overview of EF-GAIfO. Robot experience is used to estimate the feasibility of state-only demonstration transitions, and the estimated feasibility weights are incorporated into discriminator training for feasibility-aware policy learning.
Fig. 3 : Demonstration trajectories in the point-mass environment for (a) the multi-speed setting and (b) the multi-path setting. Red and blue trajectories represent infeasible and feasible demonstrations, respectively. Color intensity indicates execution speed, with lighter colors denoting slower motions.
Setting
Method
Success rate ( ↑ )
Hausdorff distance ( ↓ )
EF-GAIfO
90.0±30.0%
0.884±1.72
Multi-speed
WGAIfO
20.0±44.7%
2.629±2.694
GAIfO
15.0±35.7%
4.877±2.396
EF-GAIfO
100.0±0.0%
0.163±0.075
Multi-path
WGAIfO
20.0±41.0%
1.742±2.343
GAIfO
50.0±50.0%
2.828±2.57
TABLE I : Mean task success rates and Hausdorff distances in the point-mass experiment. The Hausdorff distance measures the discrepancy between the trajectories generated by the learned policy and the feasible demonstration trajectories. Values are reported as the mean ± standard deviation of the learned policy over 20 random seeds. Each learned policy is evaluated over ten episodes.
Setting
Demonstration
Iter.
EF-GAIfO
WGAIfO
Feasible ( ↑ )
100
0.142 ± 0.032
0.450 ± 0.039
Multi-
Feasible ( ↑ )
3000
0.924 ± 0.105
0.402 ± 0.473
speed
Infeasible ( ↓ )
100
0.016 ± 0.030
0.023 ± 0.005
Infeasible ( ↓ )
3000
0.141 ± 0.166
0.0001 ± 0.0004
Feasible ( ↑ )
100
0.152 ± 0.031
0.301 ± 0.068
Multi-
Feasible ( ↑ )
3000
0.721 ± 0.021
0.298 ± 0.030
TABLE II : Mean evaluation values for feasible and infeasible demonstration data across training iterations, assigned using EF-GAIfO and WGAIfO trained with 20 different random seeds in the point-mass experiments. The evaluation values should be high for feasible demonstrations and low for infeasible demonstrations.
Fig. 4 : Feasibility weights estimated by EF-GAIfO in the point-mass experiments. (a) and (b) show the multi-speed setting at 100 and 3000 training iterations, respectively, while (c) and (d) show the multi-path setting. The color of each demonstration trajectory represents its estimated feasibility weight.
Fig. 5 : Object-reaching-and-picking task setting in the quadruped robot environment.
Fig. 6 : Demonstration for robot experiment
Fig. 7 : Human-collected object-reaching-and-picking demonstrations used in the quadruped robot experiment. The demonstrations include variations in planar position and pitch motion.
EF-GAIfO
WGAIfO
GAIfO
Target 1
83.6±9.2%
34.8±47.7%
37.4±14.2%
Target 2
88.0±5.7%
35.2±47.8%
36.8±8.1%
Target 3
84.8±10.0%
39.0±48.8%
42.0±17.6%
Random target
85.8±7.4%
26.4±44.1%
35.4±18.1%
TABLE III : Object-reaching-and-picking task success rates in simulation for three fixed target positions and a random target setting. Each value is the mean success rate over five policies trained with different random seeds, with each policy evaluated over 100 trials.
Fig. 8 : Execution of the learned policies in the object-reaching-and-picking task using the Go2 robot in simulation. (a) EF-GAIfO successfully completes the task. (b) WGAIfO fails due to excessive pitch motion. (c) GAIfO fails due to misalignment between the gripper and the object.
Fig. 9 : Feasibility weights estimated by EF-GAIfO for demonstrated pitch motions in the quadruped robot experiment at (a) 100 and (b) 3000 training iterations.
EF-GAIfO
WGAIfO
GAIfO
Target 1
70%
10%
10%
Target 2
70%
10%
10%
Target 3
80%
10%
20%
Random target
80%
30%
20%
TABLE IV : Real-world object-reaching-and-picking task success rates for three fixed target positions and a random target setting. Each setting is evaluated over 10 trials.
Fig. 10 : Representative real-world execution examples of the learned policies in the object-reaching-and-picking task using the Go2 robot. The snapshots illustrate typical behaviors observed for each method: (a) EF-GAIfO successfully approaches, grasps, and lifts the object; (b) WGAIfO exhibits a grasping failure; and (c) GAIfO exhibits a grasping failure caused by gripper–object misalignment.
Behavior cloning for robot manipulation relies on expert demonstrations. However, for tasks that require dynamic stability, precise contact timing, or dexterous coordination, human operators may find it hard or even impossible to collect data. We study this infeasible-demonstration regime and propose GLIDE: Guardrails for Learning from Infeasible Demonstrations Efficiently, a framework that infers task-specific failure modes and converts them into executable guardrails for data collection and policy deployment. Given a task description and the conditioning teleoperation code, GLIDE writes guardrails that use system states to filter teleoperation and policy commands, constrain failure-prone actions, and iteratively improve from trajectory feedback. Across three tasks, GLIDE discovers emergent guardrails that go beyond domain-expert hardcoded ones, improving data collection over naive VR teleoperation and domain-expert hardcoded guardrails. After refinement, GLIDE raises data-collection success from 0-10 percent to 70-90 percent across the three tasks. During policy execution, mixed-data guarded policies reach 70 percent, 60 percent, and 60 percent success on Tomato plate transfer, Marker handover and stand, and Wine serving tasks. These results show that GLIDE can support policy learning when direct demonstrations are infeasible. Project website: http://guardrail-policy.github.io/
Robotic imitation learning is often treated as reproducing demonstrated actions, but actions are inherently embodiment-specific. When demonstrations come from humans or robots with different morphology, kinematics, or action spaces, this action-centric view requires shared action spaces, heuristic retargeting, or large-scale multi-embodiment co-training. We instead view demonstrations as implicit specifications of future goals: the target agent should infer what state the demonstrator is trying to realize, rather than how the demonstrator executes it. We propose Demo-JEPA, a cross-embodiment imitation framework that decouples demonstration intent from embodiment-specific execution. Built on a JEPA-based world model, Demo-JEPA translates source visual demonstrations into target-compatible future latent trajectories in a shared predictive representation space. The target agent then uses these latent trajectories as subgoals and realizes them through planning under its own learned forward dynamics. Because Demo-JEPA avoids action-level correspondence and requires only visual demonstrations plus the target agent's own interaction experience, it supports flexible imitation across heterogeneous embodiments. Experiments on RLBench and real-world manipulation tasks show that Demo-JEPA matches specialized in-domain planners and generalizes to unseen tasks and embodiment configurations where prior methods fail.
Jingyang He, Guangrun Li, Jieyu Zhang +3
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Beijing Innovation Center of Humanoid Robotics · University of Washington
Imitation learning offers a promising framework for enabling robots to acquire diverse skills from human users. However, most imitation learning algorithms assume access to high-quality demonstrations an unrealistic expectation when collecting data from non-expert users, whose demonstrations often contain inadvertent errors. Naively learning from such demonstrations can result in unsafe policy behavior, while discarding entire demonstrations due to occasional mistakes wastes valuable data, especially in low-data settings. In this work, we introduce GiB (Good-in-Bad), an algorithm that automatically identifies and discards erroneous subtasks within demonstrations while preserving high-quality subtasks. The filtered data can then be used by any policy learning algorithm to train more robust policies. GiB first trains a self-supervised model to learn latent features and assigns binary weights to label each demonstration as good or bad. It then models the latent feature distribution of high-quality segments and uses the Mahalanobis distance to detect and evaluate poor-quality subtasks. We validate GiB on the Franka robot in both simulated and real-world multi-step tasks, demonstrating improved policy performance when learning from mixed-quality human demonstrations.
Noushad Sojib, Ola Ghattas, Momotaz Begum
Department of Computer Science, University of New Hampshire, USA