Beyond State-as-Action: Exploiting Command-State Discrepancy for Robot Imitation Learning
Authors: Peiyan Li, Yueran Tao, Enhao Zhang, Zhixuan Zhao, Chenghao Yue, Hao Wang, Lei Lv, Wentao Zhao, +5 more
Organizations: Tsinghua University · SEEN·E Robotics · Imperial College London · Dalian University of Technology · Tongji University · Peking University
Constructing action targets from measured robot motion is an established approach in imitation learning. Under interaction constraints, however, command-state discrepancy may reflect control demands that motion alone does not capture. We investigate when this information matters and how to exploit it. Across three real-robot tasks, task and phase analyses reveal larger supervision gaps under constrained interaction, while selective command retention provides evidence of locally useful command information. Building on these findings, we propose Command-State Discrepancy Weighting (CSDW), which accounts for robot response times and combines subsequent progress, persistent unmet demand, and demand changes into continuous weights for command supervision. The method requires no task-phase annotations or changes to policy architecture or inference. CSDW improves over uniform command supervision on constrained tasks, while methods perform similarly in the less constrained task. Project page: https://seen-e.github.io/CSDW/.
Figures & tables
Fig. 1: Control intent in command–state discrepancy. During plug insertion, contact limits further robot motion despite continued insertion commands. The gap between the command target and the robot’s current measured state can indicate the operator’s intent to continue insertion, even when actual motion is limited. Faded poses indicate earlier configurations, and the lower images show contact followed by full insertion.
Fig. 3: Temporal evidence for held and reduced insertion requests. The blue solid and gray dashed curves show retained command c and directional state progress ρ , respectively, normalized by the initial command–state gap and clipped to [0,1] . Zero ρ does not necessarily imply a stationary robot. X and E are channel-wise unfulfilled and corrected effort scores, not measured forces; w is the final frame weight after pooling and temporal processing. Images provide context only.
Fig. 4: CSDW weights in a cable-insertion demonstration. Robot views contextualize pickup, handover, and insertion. Colored areas show the contributions of delayed progress, sustained effort, and effort changes above the unit baseline. Red shows the final weight. Additional weight extends across the annotated insertion interval, with local variation within it. Phase labels are manual annotations used only to interpret this qualitative example.
Fig. 5: Task setups and representative demonstrations for (a) cable insertion, (b) connector mating, and (c) package transfer. The first two columns show the workspace and task objects; the remaining six show successive task stages using head and wrist views. Blue labels indicate the evaluated stages: alignment and full insertion for (a) and (b), and pickup and placement for (c). Gold shading marks selected insertion stages involving sustained constrained interaction. Arrows highlight relevant objects and interaction points. Timestamps are relative to each demonstration.
Fig. 6: Robot setup and demonstration collection. (a) The Marvin robot, camera locations, and corresponding RGB views; left and right refer to the robot. (b) VR teleoperation and the recording of aligned commands, measured states, and RGB images.
Cable insertion (35)
Connector mating (30)
Package transfer (30)
Method
Alignment
Full insertion
Alignment
Full insertion
Pickup
Placement
State
77.1%
25.7%
26.7%
0.0%
96.7%
80.0%
Command
77.1%
60.0%
36.7%
33.3%
86.7%
80.0%
Hybrid
71.4%
71.4%
36.7%
36.7%
93.3%
76.7%
CSDW
94.3%
94.3%
73.3%
73.3%
100.0%
83.3%
TABLE I: Stage success rates across three tasks.
Fig. 8: Representative State and Command execution outcomes in the two insertion tasks. Close-ups show partial engagement under State supervision and full engagement under Command supervision, illustrating the distinction between alignment and final completion in Table I .
Fig. 9: Selective command retention in a connector mating demonstration. Keyframes are linked to the Hybrid label timeline: red intervals use Command arm targets and gray intervals use State targets. Retained intervals include brief grasping events and a longer period spanning approach, alignment, and insertion. Selection uses the 60th percentile of scores over the training dataset; Command targets occupy 53.5% of this demonstration’s frames.
College of Connected Computing, Vanderbilt University, USA. · School of Computer Science, University of Waterloo, Canada. · School of Computer Science, The University of Sydney, Australia. +1