Organizations: Division of Information Science, Graduate School of Information Science, Nara Institute of Science and Technology (NAIST), Nara, Japan · Department of Industrial Engineering of the University of Trento, Trento, Italy
With the increasing use of robot-free demonstration interfaces that provide state trajectories without action labels, imitation from observation has become a promising approach for learning robot behaviors from human demonstrations. However, due to differences in embodiment and dynamics between humans and robots, demonstrated human motions may not be feasible for the robot, potentially degrading policy performance. In this study, we propose Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO), which estimates the feasibility of state-only demonstrations from the robot's own experience rather than relying on explicit dynamics models or large prior exploration datasets. A key feature of EF-GAIfO is that the notion of feasibility evolves with policy learning: as the policy improves and the robot experiences a broader range of state transitions, the feasible region is progressively expanded, allowing additional demonstrations to be incorporated into learning. This enables feasibility-aware imitation that adapts to the current stage of policy learning, rather than relying on a pre-designed feasibility criterion. We validate the effectiveness of EF-GAIfO on a locomotion task in simulation and on a real quadruped robot performing a object-reaching-and-grasping task.
Figures & tables
Fig. 1 : Evolution of feasibility estimation during policy learning. Early in learning, limited robot experience yields high feasibility estimates only for demonstration transitions near the experienced region. As experience accumulates, the covered region expands, allowing more demonstration transitions to receive high feasibility estimates. This progressively reduces the influence of infeasible demonstrations during imitation learning.
Fig. 2 : Overview of EF-GAIfO. Robot experience is used to estimate the feasibility of state-only demonstration transitions, and the estimated feasibility weights are incorporated into discriminator training for feasibility-aware policy learning.
Fig. 3 : Demonstration trajectories in the point-mass environment for (a) the multi-speed setting and (b) the multi-path setting. Red and blue trajectories represent infeasible and feasible demonstrations, respectively. Color intensity indicates execution speed, with lighter colors denoting slower motions.
Setting
Method
Success rate ( ↑ )
Hausdorff distance ( ↓ )
EF-GAIfO
90.0±30.0%
0.884±1.72
Multi-speed
WGAIfO
20.0±44.7%
2.629±2.694
GAIfO
15.0±35.7%
4.877±2.396
EF-GAIfO
100.0±0.0%
0.163±0.075
Multi-path
WGAIfO
20.0±41.0%
1.742±2.343
GAIfO
50.0±50.0%
2.828±2.57
TABLE I : Mean task success rates and Hausdorff distances in the point-mass experiment. The Hausdorff distance measures the discrepancy between the trajectories generated by the learned policy and the feasible demonstration trajectories. Values are reported as the mean ± standard deviation of the learned policy over 20 random seeds. Each learned policy is evaluated over ten episodes.
Setting
Demonstration
Iter.
EF-GAIfO
WGAIfO
Feasible ( ↑ )
100
0.142 ± 0.032
0.450 ± 0.039
Multi-
Feasible ( ↑ )
3000
0.924 ± 0.105
0.402 ± 0.473
speed
Infeasible ( ↓ )
100
0.016 ± 0.030
0.023 ± 0.005
Infeasible ( ↓ )
3000
0.141 ± 0.166
0.0001 ± 0.0004
Feasible ( ↑ )
100
0.152 ± 0.031
0.301 ± 0.068
Multi-
Feasible ( ↑ )
3000
0.721 ± 0.021
0.298 ± 0.030
TABLE II : Mean evaluation values for feasible and infeasible demonstration data across training iterations, assigned using EF-GAIfO and WGAIfO trained with 20 different random seeds in the point-mass experiments. The evaluation values should be high for feasible demonstrations and low for infeasible demonstrations.
Fig. 4 : Feasibility weights estimated by EF-GAIfO in the point-mass experiments. (a) and (b) show the multi-speed setting at 100 and 3000 training iterations, respectively, while (c) and (d) show the multi-path setting. The color of each demonstration trajectory represents its estimated feasibility weight.
Fig. 5 : Object-reaching-and-picking task setting in the quadruped robot environment.
Fig. 6 : Demonstration for robot experiment
Fig. 7 : Human-collected object-reaching-and-picking demonstrations used in the quadruped robot experiment. The demonstrations include variations in planar position and pitch motion.
EF-GAIfO
WGAIfO
GAIfO
Target 1
83.6±9.2%
34.8±47.7%
37.4±14.2%
Target 2
88.0±5.7%
35.2±47.8%
36.8±8.1%
Target 3
84.8±10.0%
39.0±48.8%
42.0±17.6%
Random target
85.8±7.4%
26.4±44.1%
35.4±18.1%
TABLE III : Object-reaching-and-picking task success rates in simulation for three fixed target positions and a random target setting. Each value is the mean success rate over five policies trained with different random seeds, with each policy evaluated over 100 trials.
Fig. 8 : Execution of the learned policies in the object-reaching-and-picking task using the Go2 robot in simulation. (a) EF-GAIfO successfully completes the task. (b) WGAIfO fails due to excessive pitch motion. (c) GAIfO fails due to misalignment between the gripper and the object.
Fig. 9 : Feasibility weights estimated by EF-GAIfO for demonstrated pitch motions in the quadruped robot experiment at (a) 100 and (b) 3000 training iterations.
EF-GAIfO
WGAIfO
GAIfO
Target 1
70%
10%
10%
Target 2
70%
10%
10%
Target 3
80%
10%
20%
Random target
80%
30%
20%
TABLE IV : Real-world object-reaching-and-picking task success rates for three fixed target positions and a random target setting. Each setting is evaluated over 10 trials.
Fig. 10 : Representative real-world execution examples of the learned policies in the object-reaching-and-picking task using the Go2 robot. The snapshots illustrate typical behaviors observed for each method: (a) EF-GAIfO successfully approaches, grasps, and lifts the object; (b) WGAIfO exhibits a grasping failure; and (c) GAIfO exhibits a grasping failure caused by gripper–object misalignment.
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Beijing Innovation Center of Humanoid Robotics · University of Washington