Robot manipulation data collection has been shifting from teleoperation toward robot-free demonstrations, through interfaces such as the Universal Manipulation Interface (UMI) or directly from human hands. Vision-Language-Action (VLA) policies trained on such data inherit the demonstrator's timing. Yet human timing does not directly transfer to robots: compliant hands tolerate fast contact, whereas robots may overshoot due to actuator and tracking limitations; conversely, robots can move faster in free space. This motivates a unified approach that reconciles execution speed with contact safety. We present RoboPace, an online retiming layer that preserves the policy's geometric path while adapting its timing, respecting the target robot's kinematic and dynamic constraints. It adapts execution speed based on predicted contact, jointly accounting for contact-dependent speed limits and the robot's motion constraints. The method requires no policy retraining and operates in real time. Across three contact-rich tasks on a dual-arm robot, faster uniform execution and physical-limit-only retiming largely fail. RoboPace instead achieves higher overall success than slow uniform execution while completing four of five commands in approximately half the time, retaining the reliability of slow execution without its time cost.
Figures & tables
Figure 1: Contact-Aware Retiming of Action Chunks: RoboPace preserves the geometric path a policy predicts and adapts only its timing while respecting the robot’s kinematic and dynamic constraints: predicted contact imposes a speed limit on the hands at contact-critical waypoints (red), while free-space motion (green) runs as fast as those constraints allow. On a dual-arm robot, this achieves higher overall success than uniformly slow execution while completing four of five commands in about half its time.
Figure 2: RoboPace System Overview: TOPP-RA retimes each VLA action chunk as a joint path, limiting speed at waypoints where contact is predicted. This accelerates free-space motion while slowing at contact.
Figure 3: RoboPace Pipeline: Geometric and camera cues identify contact-critical waypoints for TOPP-RA retiming. The play-out feeds its commanded velocity into the next solve, allowing consecutive action chunks to join without stopping.
Figure 4: Contact Predictor Evaluation: (a) predicted and ground-truth stc on one Cup test clip, (b) mean absolute error (MAE) versus ground-truth stc , and (c) per-task MAE at contact onset ( stcgt=0 ). At contact onset, MAE is below 1.3 source steps for all tasks.
Figure 5: Comparison with Baselines: Accelerated execution and physical-limit-only retiming rarely succeed, whereas decelerated execution is reliable but slow. RoboPace achieves 32 successes in 50 trials and completes four of five commands in about half the decelerated execution time.
Figure 6: Execution Trace of a Bimanual Handover: One run of Part , showing that RoboPace slows as the right hand approaches the part and during hand–hand contact, while accelerating elsewhere. Red and green indicate gated and ungated motion, respectively.
vcontact (m/s)
schedule
0.01
0.02
0.04
Linear- N3
0.6 (23.1)
0.8 (16.1)
0.4 (10.0)
Linear- N9
1.0 (26.3)
0.2 (13.1)
0.0 (–)
Cosine- N3
0.8 (25.4)
0.4 (15.6)
0.2 (8.0)
Cosine- N9
0.6 (23.8)
0.8 (22.4)
0.2 (8.0)
Tanh- N3
1.0 (32.5)
0.8 (16.9)
0.2 (9.0)
Table 1: Sensitivity to the contact speed limit and transition schedule. (a) Five trials per setting on Pyramid phase 2: success is more sensitive to the contact speed limit, with lower limits trading speed for reliability. (b) Ten trials per command across five commands: Linear- N3 (bold) achieves the highest overall success ( 32/50 ) among the three lowest-cost schedules selected at 0.02 m/s and is used as the default.
Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VLAs) only inherit a single fixed speed from training demonstrations. Prior efforts to accelerate VLAs through model compression, KV-cache reuse, or reinforcement learning only shift the policy from one fixed speed to another, and leave deceleration almost unexplored. We observe that the magnitude of each predicted action already governs how fast the robot moves, opening a direct route to controllable execution speed. We turn this observation into TempoVLA, a single VLA whose execution speed is controlled by an explicit condition. TempoVLA combines two coupled components. (1) A data-side Variable-Speed Trajectory Augmentation (VSTA) that re-times demonstration to any target speed by merging or splitting actions while preserving its motion semantics. (2) A model-side conditioning mechanism that feeds the speed to the policy. Statistics show that VSTA reaches the requested speed with negligible motion error. Experiments in simulation and on real-world tasks demonstrate that TempoVLA achieves flexible speed control in both directions, while VSTA additionally boosts the default 1× performance via better data utilization. Furthermore, by cooperating with a large multimodal model, TempoVLA realizes dynamic speed control, accelerating through low-risk phases and decelerating for high-risk ones.
Recent vision-language-action and diffusion-based robot policies often use action chunking, where each policy query predicts a sequence of future actions and the robot executes an open-loop prefix before re-querying. While this interface improves local motion continuity, deployment still requires choosing the execution horizon: how much of each predicted chunk should be executed before acquiring a new observation. However, our experiments show that success is strongly task-dependent and non-monotonic with respect to the execution horizon, making a single constant horizon an unreliable deployment rule. We propose PACE (Phase-Aware Chunk Execution), a training-free test-time execution method that selects the execution horizon online from the predicted chunk itself. PACE exploits the phase-dependent kinematic structure of manipulation trajectories by identifying low-speed transition points in the predicted speed profile and using them as candidate replanning boundaries. Because PACE uses only the predicted action chunk, it is plug-and-play and requires no retraining or access to policy internals. We validate PACE through large-scale evaluations in both simulation and real-robot settings. On 50 RoboTwin2.0 tasks, PACE raises the average success rate from 57.8% to 64.2%. In real-robot experiments on bimanual ALOHA and single-arm Franka platforms, PACE improves the average task score from 60.7 to 77.7 and the average success rate from 50.7% to 70.4%. Ablations and rollout-level analyses show that PACE adapts execution horizons across manipulation phases, shortening near transitions while preserving longer execution during coherent motion.
Junnan Nie, Jiayi Li, Chenghao Liu +5
Peking University · Peking University. · JD Explore Academy. +1
Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce EgoSpeedUp, a framework that uses human manipulation as temporal supervision for robot imitation learning. Our key insight is that human demonstrations naturally reveal task-appropriate, phase-wise manipulation tempo. Given slow robot demonstrations and human demonstrations of the same task, EgoSpeedUp aligns corresponding manipulation phases, estimates their relative execution tempos from multiple human demonstrations, and transfers the resulting phase-wise tempo by retiming the robot demonstrations. The retimed demonstrations are then used for standard behavior cloning, allowing the robot to retain its executable manipulation behavior while learning to perform it at a human-informed tempo. Across two real-world manipulation tasks, EgoSpeedUp improves the task success rate by an average of 25 percentage points (pp) while reducing successful execution time by 36.5%. These results demonstrate that human manipulation tempo provides an effective temporal reference for learning faster and more reliable robot policies.
Hanbit Oh, Yukiyasu Domae, Takuma Yagi
Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), Japan