We introduce Task-Space Imitation Guidance for Efficient Reinforcement Learning (TIGER), a reward-construction and pretraining framework for sparse-reward tabletop robotic manipulation. TIGER treats an action-chunked imitation policy not as an executable controller or action prior, but as a local task-space progress estimator: predicted action chunks are converted, using controller-aware action-to-motion mapping, into short-horizon end-effector references, and the RL agent receives dense progress rewards toward these references while the sparse environment reward remains the dominant objective. During pretraining, TIGER uses imitation-guided look-ahead signals to relax conservative value penalties for actions predicted to make task-space progress, reducing off-manifold exploration during early online RL. Across simulation and real-robot experiments, TIGER improves early sample efficiency and reduces measured safety violations while matching or improving final success rates relative to prior RL and IL-RL baselines on the evaluated tasks.
Figures & tables
Robomimic
MetaWorld
Humanoid
Method
Can
Square
ToolHang
Assembly
Box Close
Stick Pull
Shape Sorting
BC
–/–/0.7
–/–/0.6
–/–/0.5
–/–/0.7
–/–/0.38
–/–/0.25
–/–/0.1
LaNE
–/–/0.25
–/–/0.025
–/–/0.0
–/–/0.08
–/–/0.06
–/–/0.6
–/–/0.0
RLPD
55/–/0.71
120/–/0.79
–/–/0.0
40/50/0.95
35/45/0.94
–/–/0.18
–/–/0.0
IBRL
35/75/0.91
80/160/0.91
470/–/0.0
25/30/0.96
25/40/0.94
35/–/0.72
115/–/0.78
DSRL
5 /–/0.64
–/–/0.37
–/–/0.0
5 /–/0.73
–/–/0.3
–/–/0.12
–/–/0.0
Table 1: Simulation benchmark summary. Each cell reports S50↓/S90↓/F↑ , where S50 and S90 are environment steps, in thousands, required to reach 50% and 90% success, and F is final success. “–” denotes that the threshold was not reached.
Figure 1: Final episode length achieved by each method.
Task
ReWiND
Robometer
TIGER
Assembly
25/30/ .95
25/30/.95
25/30/.95
Box Close
30/60/.87
25/50/.92
25/45/.93
Stick Pull
50/60/.84
55/–/.60
40/50/.92
Square
130/–/.74
120/–/.71
90/160/.87
Table 2: Additional simulation comparisons. Entries report S50↓/S90↓/F↑ ; environment steps are in thousands and results are averaged over five seeds. Panel (a) uses the same demonstrations and RLPD backend. Panel (b) reports task averages over all ten LIBERO-Spatial tasks; BC has no online-learning thresholds. A dash otherwise indicates that a threshold was not reached.
Figure 2: Top: Visualization of each task. Bottom: Training curves for each task. The x-axis represents environment steps in intervals of 1000. For block picking and drawer opening, the y-axis indicates the average task success rate over each 1000-step interval. For rotary insertion, the y-axis denotes the curriculum stage. Task completion requires reaching stage #9.
Figure 3: Safety violations on the two real-robot tasks (block picking and drawer opening). Bars show the mean number of episode resets across two seeds with min–max ranges.
Stationary
Rotating
Random Location
Location
BC
60%
10%
TIGER
100%
90%
IBRL
90%
77.5%
DSRL
60%
20%
Table 3: Rotary insertion evaluation.
Figure 7
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Performance on MetaWorld benchmark
Figure 7: Performance on Robomimic benchmark
Figure 8: Performance on Humanoid Shape Sorting task
Figure 11: Effect of demonstration quality on learning performance
Figure 12: Effect of the reward-scaling factor k on success rate.
Figure 13: Pretraining objective ablation on Robomimic Can.
Figure 14: Transferability of the controller-aware mapping.
Figure 15: Sensitivity of controller-aware mapping accuracy to the number of training demonstrations. Validation MSE decreases rapidly and saturates after approximately five demonstrations, indicating that the 50-demonstration fits used in production are not data-limited.
Long-horizon policies trained with reinforcement learning can still achieve high return through inefficient interactions, while rare efficient behaviors discovered during training may be forgotten. We argue that temporal efficiency itself provides a source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 30 long-horizon tasks spanning robot manipulation and interactive navigation, TSIL consistently improves task success rates, learning efficiency, behavioral efficiency, and robustness to unstable training conditions.
Learning robot manipulation policies with deep neural networks from a single demonstration remains highly challenging, as even small deviations from the demonstrated trajectory can quickly compound into failure, while collecting substantial online interaction data is costly. We propose ReGIL, a retrieval-guided imitation learning framework that treats a single demonstration as an external memory. ReGIL repeatedly queries this static memory throughout training to simultaneously guide exploration, generate the regularization buffer, and construct rewards. Specifically, it computes rewards through local temporal alignment between the current trajectory and the retrieved segment, providing step-wise and informative feedback for policy improvement. We evaluate ReGIL on robotic manipulation tasks from the LIBERO and Meta-World benchmarks under the single demonstration setting. ReGIL outperforms prior baselines in both success rate and training efficiency. In real-robot experiments, using only one demonstration and less than one hour of online training, ReGIL achieves over 75% success rate across three manipulation tasks with randomness in both initial robot pose and target position. These results demonstrate that leveraging the single demonstration as reusable memory can provide more than static supervision for efficient robot learning. More details can be found on our website: https://regil2026.github.io/
Yuying Zhang, Francesco Verdoja, Wenyan Yang +1
School of Electrical Engineering Aalto University, Finland
While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by suboptimal execution speeds. Imitation learning policies are inherently limited by hardware constraints and the speed of the operator during data collection. In addition, there are no established methods for accelerating policies learned via imitation, and the empirical relationship between execution speed and task success remains underexplored. To address these issues, we introduce SpeedTuning, a reinforcement learning framework specifically designed to enhance the speed of manipulation policies. SpeedTuning learns to predict the optimal execution speed for actions, thereby complementing a base policy without necessitating additional data collection. We provide empirical evidence that SpeedTuning achieves substantial improvements in execution speed, exceeding 2.4x speed-up, while preserving an adequate success rate compared to both the original task policy and straightforward speed-up methods such as linear interpolation at a fixed speed. We evaluate our approach across a diverse set of dynamic and precise tasks, including pouring, throwing, and picking, demonstrating its effectiveness and robustness in enhancing real-world robotic manipulation. Videos and code are available at https://daivdyuan.github.io/speed-tuning/