cs.ROOct 7, 2026

TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies

Authors: Yeonseo Lee, Hyosup Shin, Guebin Hwang, Sungho Jo

Organizations: Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea.

Abstract

Vision-Language-Action (VLA) models are effective at understanding what task to perform, but provide limited control over how it should be executed, such as moving quickly or slowly. We introduce TempoBridge, a lightweight framework that uses frozen VLA representations to modulate actions according to tempo cues in the instruction at each task phase, without additional tempo-conditioned robot demonstrations or tempo-specific base-policy fine-tuning. TempoBridge extracts tempo cues from contextual VLM representations, aligns them with task progress through a causal phase router, and modulates nominal motion commands during execution. Across LIBERO tasks, TempoBridge improves Tempo Success Rate from 52.6% to 89.7% under canonical tempo instructions while retaining high task success. It also preserves near-baseline performance when no tempo cue is present and generalizes to unseen tempo expressions without additional training. Experiments on a physical robot further demonstrate language-conditioned tempo modulation in real-world manipulation.

Figures & tables

Explore similar work

CardsList
  1. TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

    Jun 4, 2026Dong Jing, Jingchen Nie, Tianqi Zhang +4

  2. Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

    Jul 14, 2026Yuzhou Wu, Yuxin Zheng, Muchun Niu +6Diffusion-Based Vision-Language-ActionsVision-Language-Action Framework

  3. Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference

    May 4, 2026Yudong Liu, Yuan Li, Zijia Tang +12Robotic ManipulationLatent Variable