Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace. This coupling complicates adaptation from teleoperation: useful geometric guidance comes with timing shaped by interface delays and operator behavior. Our key insight is to bring the path-time parameterization of classical motion planning into the learned action representation of a VLA. We introduce PathTime-VLA, which represents motion as a progress-indexed interaction path X(s) and a positive interval-time profile. The latter defines a monotone time law t(s), yielding controller commands X(s(t)). For a given path, alternative executions are expressed through the time profile, allowing chunk-wise speed choices without changing the geometric prediction target. This representation supports a staged post-training procedure: demonstrations and DAgger interventions establish a target-domain prior, Speed-DQN learns execution multipliers from robot interaction, and Path-AWR uses rollout outcomes to refine the diffusion path generator. A path-conditioned action expert realizes the resulting motions while maintaining distinct learning interfaces for path generation and execution timing. Across three tasks, the complete method achieves 58/60 successes versus 57/60 for PathTime-VLA under BC + DAgger at fixed 1×, with approximately 39-52% shorter mean completion times over successful trials.
Figures & tables
Fig. 1: Path-time parameterization for VLA actions. Our representation separates fixed-progress path samples from interval durations defining t(s) . Bottom: chunk-wise replanning and aggregate real-robot results for Coupled + DAgger versus the complete PathTime-VLA pipeline; throughput is averaged across tasks.
Fig. 2: Path-Time action expert. The PathHead predicts fixed-progress states; the rate and timing branches determine their interval durations.
Fig. 3: Factorized post-training. The path-time representation exposes alternative timing choices for a shared path target. Supervised adaptation establishes the action expert; Speed-DQN explores these choices on the robot; Path-AWR uses the resulting experience to refine path generation.
Fig. 4: Real-world platform and task settings. All tasks use a Franka Panda, two RGB views, matched randomized initial conditions, and a 60 -s horizon.
Peg insertion
Microwave opening
Tape + drawer
All tasks
Training
Representation
S
T
S
T
S
T
S
BC
Coupled
9/20
18.00
14/20
33.43
17/20
27.94
40/60
PathTime-VLA
9/20
14.33
15/20
30.27
19/20
26.15
43/60
BC + DAgger
Coupled
16/20
31.68
16/20
32.94
19/20
26.68
51/60
PathTime-VLA
19/20
17.68
19/20
32.05
19/20
23.37
57/60
TABLE I: Comparison of action representations. S : successes out of 20; T : mean successful time (s).
Fig. 5: Success under fixed execution multipliers with BC and BC + DAgger. Labels report absolute success rates; whiskers show 95% Wilson intervals. Highlighted labels indicate whether a setting performs above (blue) or at/below (orange) that policy’s own 1× success rate.
Fig. 6: Task-wise post-training success and throughput. Top: success rate with counts out of 20 and 95% Wilson intervals. Bottom: throughput proxy with task-specific scales. BC and BC + DAgger use fixed 1× ; AWR retains the DQN scheduler.
Fig. 7: Execution progress for PathTime-VLA under BC + DAgger at fixed 1× and the complete method. Rows show representative rollouts; checks and crosses denote success and failure, and arrows mark the shown rollout durations.