Vision-Language-Action (VLA) models are effective at understanding what task to perform, but provide limited control over how it should be executed, such as moving quickly or slowly. We introduce TempoBridge, a lightweight framework that uses frozen VLA representations to modulate actions according to tempo cues in the instruction at each task phase, without additional tempo-conditioned robot demonstrations or tempo-specific base-policy fine-tuning. TempoBridge extracts tempo cues from contextual VLM representations, aligns them with task progress through a causal phase router, and modulates nominal motion commands during execution. Across LIBERO tasks, TempoBridge improves Tempo Success Rate from 52.6% to 89.7% under canonical tempo instructions while retaining high task success. It also preserves near-baseline performance when no tempo cue is present and generalizes to unseen tempo expressions without additional training. Experiments on a physical robot further demonstrate language-conditioned tempo modulation in real-world manipulation.
Figures & tables
Fig. 2: Overview of TempoBridge. Using the frozen VLM backbone within π0.5 , a prototype-based tempo readout extracts an ordered tempo sequence from the instruction once and caches it for execution. At each replan, a causal phase router uses visual features reused from the policy forward pass and robot state to track task progress and select the active event. The corresponding tempo coefficient scales the translational components of the nominal action chunk.
Method
Spatial
Object
Goal
Long
Overall
Task Success Rate ↑
Tempo Success Rate ↑
Task Success Rate ↑
Tempo Success Rate ↑
Task Success Rate ↑
Tempo Success Rate ↑
Task Success Rate ↑
Tempo Success Rate ↑
Task Success Rate ↑
Tempo Success Rate ↑
Task & Tempo Success Rate ↑
π0.5
94.0
48.9
96.0
49.5
93.0
57.8
91.5
54.4
93.6
52.6
48.8
TempoBridge
94.5
91.0
97.5
90.8
92.5
93.8
87.0
82.9
92.9
89.7
82.0
TABLE I: Task, tempo, and task & tempo success rates (%) under quickly / slowly instructions.
Fig. 3: Qualitative examples of TempoBridge execution on LIBERO. The left two panels use quickly / slowly cues, while the right three use tempo expressions unseen during readout construction. Each panel shows the instruction and executed TCP trajectory, colored by measured TCP speed from slower motion (blue) to faster motion (red).
Method
Spatial
Object
Goal
Long
Overall
π0.5
95.0
97.0
95.0
93.0
95.0
TempoBridge
98.0
96.0
95.0
88.0
94.3
TABLE II: Task success rates (%) on LIBERO benchmark instructions without tempo cues.
Method
Spatial
Object
Goal
Long
Overall
Task Success Rate ↑
Tempo Success Rate ↑
Task Success Rate ↑
Tempo Success Rate ↑
Task Success Rate ↑
Tempo Success Rate ↑
Task Success Rate ↑
Tempo Success Rate ↑
Task Success Rate ↑
Tempo Success Rate ↑
Task & Tempo Success Rate ↑
π0.5
95.6
50.3
99.3
49.5
96.1
61.2
89.5
53.1
95.1
53.5
50.6
TempoBridge
94.8
79.7
96.1
75.2
93.8
91.0
84.8
74.2
92.3
80.1
72.9
TABLE III: Task, tempo, and task & tempo success rates (%) for tempo expressions unseen during readout construction.
Tempo expression
Target tempo
Accuracy ↑
rapidly
FAST
98.3
swiftly
FAST
98.3
carefully
SLOW
57.9
at a slower pace
SLOW
95.0
TABLE IV: Cue-specific tempo recognition accuracy (%) for expressions unseen during readout construction.
Setting
Readout Acc. ↑
Balanced Phase Agreement ↑
Median Speed Ratio
Canonical
97.5%
89.7%
1.58 ×
Unseen
78.4%
90.4%
1.52 ×
TABLE V: Component-wise diagnostics of TempoBridge under canonical and unseen tempo expressions.
Fig. 4: Real-world experimental setup with an xArm6 robot, external and wrist-mounted Azure Kinect cameras, and the three manipulation task configurations.
Fig. 5: Real-world TCP speed under Fast → Slow (FS) and Slow → Fast (SF) instructions. Columns correspond to the three tasks, with the task-adapted π0.5 baseline in the top row and TempoBridge in the bottom row. Thin lines represent individual task-successful rollouts, connecting the mean TCP speeds of Event 1 and Event 2. Bold lines indicate the median, and error bars show the interquartile range over 10 successful rollouts per condition.
Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VLAs) only inherit a single fixed speed from training demonstrations. Prior efforts to accelerate VLAs through model compression, KV-cache reuse, or reinforcement learning only shift the policy from one fixed speed to another, and leave deceleration almost unexplored. We observe that the magnitude of each predicted action already governs how fast the robot moves, opening a direct route to controllable execution speed. We turn this observation into TempoVLA, a single VLA whose execution speed is controlled by an explicit condition. TempoVLA combines two coupled components. (1) A data-side Variable-Speed Trajectory Augmentation (VSTA) that re-times demonstration to any target speed by merging or splitting actions while preserving its motion semantics. (2) A model-side conditioning mechanism that feeds the speed to the policy. Statistics show that VSTA reaches the requested speed with negligible motion error. Experiments in simulation and on real-world tasks demonstrate that TempoVLA achieves flexible speed control in both directions, while VSTA additionally boosts the default 1× performance via better data utilization. Furthermore, by cooperating with a large multimodal model, TempoVLA realizes dynamic speed control, accelerating through low-risk phases and decelerating for high-risk ones.
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2
Dual-system Vision-Language-Action (VLA) models achieve state-of-the-art robotic manipulation but are bottlenecked by the VLM backbone, which must execute at every control step while producing temporally redundant features. We propose Latent Bridge, a lightweight model that predicts VLM output deltas between timesteps, enabling the action head to operate on predicted outputs while the expensive VLM backbone is called only periodically. We instantiate Latent Bridge on two architecturally distinct VLAs: GR00T-N1.6 (feature-space bridge) and π0.5 (KV-cache bridge), demonstrating that the approach generalizes across VLA designs. Our task-agnostic DAgger training pipeline transfers across benchmarks without modification. Across four LIBERO suites, 24 RoboCasa kitchen tasks, and the ALOHA sim transfer-cube task, Latent Bridge achieves 95-100% performance retention while reducing VLM calls by 50-75%, yielding 1.65-1.73x net per-episode speedup.
Yudong Liu, Yuan Li, Zijia Tang +12
1Duke University · 2Qualcomm AI Research · University of Florida