Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace. This coupling complicates adaptation from teleoperation: useful geometric guidance comes with timing shaped by interface delays and operator behavior. Our key insight is to bring the path-time parameterization of classical motion planning into the learned action representation of a VLA. We introduce PathTime-VLA, which represents motion as a progress-indexed interaction path X(s) and a positive interval-time profile. The latter defines a monotone time law t(s), yielding controller commands X(s(t)). For a given path, alternative executions are expressed through the time profile, allowing chunk-wise speed choices without changing the geometric prediction target. This representation supports a staged post-training procedure: demonstrations and DAgger interventions establish a target-domain prior, Speed-DQN learns execution multipliers from robot interaction, and Path-AWR uses rollout outcomes to refine the diffusion path generator. A path-conditioned action expert realizes the resulting motions while maintaining distinct learning interfaces for path generation and execution timing. Across three tasks, the complete method achieves 58/60 successes versus 57/60 for PathTime-VLA under BC + DAgger at fixed 1×, with approximately 39-52% shorter mean completion times over successful trials.
Figures & tables
Fig. 1: Path-time parameterization for VLA actions. Our representation separates fixed-progress path samples from interval durations defining t(s) . Bottom: chunk-wise replanning and aggregate real-robot results for Coupled + DAgger versus the complete PathTime-VLA pipeline; throughput is averaged across tasks.
Fig. 2: Path-Time action expert. The PathHead predicts fixed-progress states; the rate and timing branches determine their interval durations.
Fig. 3: Factorized post-training. The path-time representation exposes alternative timing choices for a shared path target. Supervised adaptation establishes the action expert; Speed-DQN explores these choices on the robot; Path-AWR uses the resulting experience to refine path generation.
Fig. 4: Real-world platform and task settings. All tasks use a Franka Panda, two RGB views, matched randomized initial conditions, and a 60 -s horizon.
Peg insertion
Microwave opening
Tape + drawer
All tasks
Training
Representation
S
T
S
T
S
T
S
BC
Coupled
9/20
18.00
14/20
33.43
17/20
27.94
40/60
PathTime-VLA
9/20
14.33
15/20
30.27
19/20
26.15
43/60
BC + DAgger
Coupled
16/20
31.68
16/20
32.94
19/20
26.68
51/60
PathTime-VLA
19/20
17.68
19/20
32.05
19/20
23.37
57/60
TABLE I: Comparison of action representations. S : successes out of 20; T : mean successful time (s).
Fig. 5: Success under fixed execution multipliers with BC and BC + DAgger. Labels report absolute success rates; whiskers show 95% Wilson intervals. Highlighted labels indicate whether a setting performs above (blue) or at/below (orange) that policy’s own 1× success rate.
Fig. 6: Task-wise post-training success and throughput. Top: success rate with counts out of 20 and 95% Wilson intervals. Bottom: throughput proxy with task-specific scales. BC and BC + DAgger use fixed 1× ; AWR retains the DQN scheduler.
Fig. 7: Execution progress for PathTime-VLA under BC + DAgger at fixed 1× and the complete method. Rows show representative rollouts; checks and crosses denote success and failure, and arrows mark the shown rollout durations.
Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VLAs) only inherit a single fixed speed from training demonstrations. Prior efforts to accelerate VLAs through model compression, KV-cache reuse, or reinforcement learning only shift the policy from one fixed speed to another, and leave deceleration almost unexplored. We observe that the magnitude of each predicted action already governs how fast the robot moves, opening a direct route to controllable execution speed. We turn this observation into TempoVLA, a single VLA whose execution speed is controlled by an explicit condition. TempoVLA combines two coupled components. (1) A data-side Variable-Speed Trajectory Augmentation (VSTA) that re-times demonstration to any target speed by merging or splitting actions while preserving its motion semantics. (2) A model-side conditioning mechanism that feeds the speed to the policy. Statistics show that VSTA reaches the requested speed with negligible motion error. Experiments in simulation and on real-world tasks demonstrate that TempoVLA achieves flexible speed control in both directions, while VSTA additionally boosts the default 1× performance via better data utilization. Furthermore, by cooperating with a large multimodal model, TempoVLA realizes dynamic speed control, accelerating through low-risk phases and decelerating for high-risk ones.
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2
Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time robotic manipulation, especially when large vision-language backbones and iterative action heads run at every control step. Existing VLA acceleration methods often optimize individual components or rely on fixed acceleration rules, treating different control steps with largely fixed computation and overlooking the non-uniform reasoning demands of sequential embodied control. Inspired by human motor control, where cognitive and feedback resources concentrate on goal-sensitive stages, we argue that VLA models should learn when to invest full computation and when to reuse prior computation. We propose ElegantVLA, a plug-in phase-adaptive inference framework that accelerates VLA models through intra-model dynamic compute scheduling. ElegantVLA introduces a lightweight scheduler that observes temporal representation similarity, robot-motion cues, and episode progress to jointly allocate computation across the vision encoder, LLM, and action head. For perception-language reasoning, the scheduler selects a five-level Vision-LLM compute mode, from full recomputation to multi-step temporal reuse, based on visual-language representation stability. For action generation, it selects a three-level denoising mode, reusing intermediate denoising states during stable motion while preserving full refinement for goal-sensitive stages. By coordinating these decisions, ElegantVLA offers a general acceleration framework for modern VLA pipelines with explicit action-generation modules, without modifying or retraining the base model. Experiments on GR00T and CogACT achieve up to 2.55x and 3.77x speedup, and on six real-world GR00T tasks ElegantVLA cuts computation by 2.18x while raising control frequency from 13.8 Hz to 26.3 Hz.
Ye Li, Huanan Liu, Kangye Ji +7
Tsinghua University · University of Illinois at Urbana-Champaign