We introduce Task-Space Imitation Guidance for Efficient Reinforcement Learning (TIGER), a reward-construction and pretraining framework for sparse-reward tabletop robotic manipulation. TIGER treats an action-chunked imitation policy not as an executable controller or action prior, but as a local task-space progress estimator: predicted action chunks are converted, using controller-aware action-to-motion mapping, into short-horizon end-effector references, and the RL agent receives dense progress rewards toward these references while the sparse environment reward remains the dominant objective. During pretraining, TIGER uses imitation-guided look-ahead signals to relax conservative value penalties for actions predicted to make task-space progress, reducing off-manifold exploration during early online RL. Across simulation and real-robot experiments, TIGER improves early sample efficiency and reduces measured safety violations while matching or improving final success rates relative to prior RL and IL-RL baselines on the evaluated tasks.
Figures & tables
Robomimic
MetaWorld
Humanoid
Method
Can
Square
ToolHang
Assembly
Box Close
Stick Pull
Shape Sorting
BC
–/–/0.7
–/–/0.6
–/–/0.5
–/–/0.7
–/–/0.38
–/–/0.25
–/–/0.1
LaNE
–/–/0.25
–/–/0.025
–/–/0.0
–/–/0.08
–/–/0.06
–/–/0.6
–/–/0.0
RLPD
55/–/0.71
120/–/0.79
–/–/0.0
40/50/0.95
35/45/0.94
–/–/0.18
–/–/0.0
IBRL
35/75/0.91
80/160/0.91
470/–/0.0
25/30/0.96
25/40/0.94
35/–/0.72
115/–/0.78
DSRL
5 /–/0.64
–/–/0.37
–/–/0.0
5 /–/0.73
–/–/0.3
–/–/0.12
–/–/0.0
Table 1: Simulation benchmark summary. Each cell reports S50↓/S90↓/F↑ , where S50 and S90 are environment steps, in thousands, required to reach 50% and 90% success, and F is final success. “–” denotes that the threshold was not reached.
Figure 1: Final episode length achieved by each method.
Task
ReWiND
Robometer
TIGER
Assembly
25/30/ .95
25/30/.95
25/30/.95
Box Close
30/60/.87
25/50/.92
25/45/.93
Stick Pull
50/60/.84
55/–/.60
40/50/.92
Square
130/–/.74
120/–/.71
90/160/.87
Table 2: Additional simulation comparisons. Entries report S50↓/S90↓/F↑ ; environment steps are in thousands and results are averaged over five seeds. Panel (a) uses the same demonstrations and RLPD backend. Panel (b) reports task averages over all ten LIBERO-Spatial tasks; BC has no online-learning thresholds. A dash otherwise indicates that a threshold was not reached.
Figure 2: Top: Visualization of each task. Bottom: Training curves for each task. The x-axis represents environment steps in intervals of 1000. For block picking and drawer opening, the y-axis indicates the average task success rate over each 1000-step interval. For rotary insertion, the y-axis denotes the curriculum stage. Task completion requires reaching stage #9.
Figure 3: Safety violations on the two real-robot tasks (block picking and drawer opening). Bars show the mean number of episode resets across two seeds with min–max ranges.
Stationary
Rotating
Random Location
Location
BC
60%
10%
TIGER
100%
90%
IBRL
90%
77.5%
DSRL
60%
20%
Table 3: Rotary insertion evaluation.
Figure 7
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Performance on MetaWorld benchmark
Figure 7: Performance on Robomimic benchmark
Figure 8: Performance on Humanoid Shape Sorting task
Figure 11: Effect of demonstration quality on learning performance
Figure 12: Effect of the reward-scaling factor k on success rate.
Figure 13: Pretraining objective ablation on Robomimic Can.
Figure 14: Transferability of the controller-aware mapping.
Figure 15: Sensitivity of controller-aware mapping accuracy to the number of training demonstrations. Validation MSE decreases rapidly and saturates after approximately five demonstrations, indicating that the 50-demonstration fits used in production are not data-limited.