Online control of execution speed is essential for deploying robot policies in real-world scenarios, as robots may need to speed up under time constraints or slow down to facilitate human interaction and improve safety. However, imitation-learned policies inherit the execution speed of their demonstrations, and test-time speed modification can introduce unrecoverable out-of-distribution observations, reducing task success. We observe that the directional inconsistency of action chunks reflects task-phase criticality, indicating how aggressively action step lengths can be modified while preserving task success. Based on this observation, we introduce FreeSpeed, a training-free module that post-processes action chunks from pretrained policies. FreeSpeed resamples each predicted chunk at the requested rate, then uses directional inconsistency between adjacent actions as the primary signal for rescaling. This signal adaptively determines how closely the execution speed can approach the requested speed, allowing flexible speed adjustment within the evaluated limits without compromising task success. Across three policy families and 50 simulated tasks, FreeSpeed supports online speed changes, with realized execution rates spanning 0.22x to 2.53x among settings that preserve per-task success. Across four real-world manipulation tasks, FreeSpeed achieves an average success rate of 94.0%, matching the frozen policy's 93.8%, while realizing execution rates from 0.38x to 1.97x.
Figures & tables
Figure 1: FreeSpeed modulates the execution speed of a frozen policy. (a) Methods based on demonstration retiming resample demonstrations and train or fine-tune the policy, while FreeSpeed resamples and rescales predicted action chunks at test time. (b) Success rates of FreeSpeed and vanilla resampling on three policies under speedup and slowdown commands.
No rate
Unmodified
Unchanged
User-set
Slow-
In-episode
training
demos
controller
rate
motion
rate
SuP ( Wu et al., 2026 )
∘
∘
∘
∘
∘
∘
SpeedTuning ( Yuan et al., 2025 )
∘
∙
∙
∘
∘
∘
TempoVLA ( Jing et al., 2026 )
∘
∘
∙
∙
∙
∙
SAIL ( Arachchige et al., 2025 )
∘
∘
∘
∘
∘
∘
RACE ( Kim et al., 2026 )
∙
∘
∘
∘
∘
∘
Table 1: Speed-control capabilities of the compared methods. A filled dot indicates that the capability is demonstrated in the corresponding paper; an open dot indicates that it is not demonstrated.
Figure 2: FreeSpeed at one replan. The frozen policy predicts an action chunk. Its translation increments provide a cosine inconsistency score, and the rate command determines temporal resampling. Together, the score and command determine the factor used to rescale the resampled translation and rotation increments.
Figure 3: Evaluation environments. 50 simulation tasks from LIBERO and four real-robot tasks.
Figure 4: The signal and the response. (a) Cosine inconsistency along human demonstrations of five LIBERO-90 tasks, with peaks during grasping and placement. (b) The mapping from cosine inconsistency to the rescaling factor under the 2× and 0.5× commands.
Type
Task
1×
1.5×
2×
0.5×
0.3×
Ref.
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
Structure constrained
2
96
72
82
94
28
36
88
74
92
46
90
32
78
76
66
76
52
50
76
58
70
66
76
40
90
70
90
88
78
78
88
100
94
90
96
86
92
70
84
92
40
64
88
82
90
86
86
Bounded receptacle
18
98
92
96
98
82
96
98
92
98
96
96
Table 2: Fixed-stride execution on 10 LIBERO-90 tasks with Flow Matching. Success rates are reported in %, with task indices defined in Table A2 . ‘Rate’ is relative to the 1× reference and averages per-task realized rates, each computed over trials in which both the evaluated method and the reference succeed. Bold indicates the highest success rate under each command. Completion steps are reported in Table A5 .
1×
2×
3×
4×
0.5×
0.3×
0.2×
Ref.
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
Peach on Plate
100
70
90
100
50
55
100
45
0
95
70
100
65
100
20
95
Tennis in Can
90
90
60
95
45
30
90
0
0
85
80
95
70
90
65
95
Stack cube
90
60
75
95
0
20
90
5
0
90
95
90
65
95
75
75
Pour almond
95
70
85
100
55
50
100
50
45
95
85
100
65
90
60
95
Mean
93.8
72.5
77.5
97.5
37.5
38.8
95.0
25.0
11.2
91.2
82.5
96.2
66.2
93.8
55.0
90.0
Table 3: Fixed-stride execution on four real-robot tasks. Per-task success rates are reported in %. Bold indicates the highest success rate under each command. Per-task completion steps and realized rates are reported in Table A10 .
Figure 5: Simulation comparison for π0.5 and Fast-WAM. Bars show success rates on the left axes, and triangular markers show realized execution rates on the right axes. The rate axes are inverted in the slowdown panels.
Figure 6: Observation distribution shift under the 2× command. FreeSpeed remains close to the 1× reference in (a) and (b), while the baselines show larger deviations from the training data, particularly after executing chunks with high cosine inconsistency in (c).
A: the stride varies
B: the factor varies
FM
π0.5
FM
π0.5
dynamic
dynamic
0.3×
0.5×
1.5×
2×
0.3×
0.5×
1.5×
2×
FreeSpeed
Succ.
–
–
90
90
90
89
97
98
97
97
Rate
–
–
0.40
0.63
1.22
1.35
0.39
0.57
1.21
1.34
varying
Succ.
90
97
89
91
92
91
97
98
96
97
Rate
0.63
0.54
0.46
0.68
1.20
1.28
0.42
0.61
1.13
1.21
Table 4: Within-episode stride and factor randomization. FM is evaluated on the 10 tasks in Table 2 , and π0.5 on all 40 tasks in Figure 5 . (A) A new stride is sampled at each replan. (B) The stride remains fixed, while the rescaling factor is sampled between fc and ft .
Figure 7: Comparison with published methods and sensitivity to the two constants. (a) Success change against realized rate. (b, c) The decay rate λ and the factor fs swept in turn.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: The real-robot workstation. The experiments use the robot’s left arm, shown on the right in the image. This frame is taken from a 1× rollout of the tennis-ball task, with the ball in the gripper and two cans on the table.
Item
Value
Robot
Rokae Helios series humanoid
Arm
7 DoF
End effector
Daimon gripper
robot command rate
30 Hz
Action interface
Cartesian end-effector increments
Head camera
Orbbec, 640×480 , 30 fps
Appendix
Table A1: Real-robot hardware and demonstration collection.
Type
Task
Language instruction
Structure constrained
2
put the black bowl in the top drawer of the cabinet
32
put the ketchup in the top drawer of the cabinet
40
put the frying pan on the cabinet shelf
86
pick up the book in the middle and place it on the cabinet shelf
Bounded receptacle
18
put the frying pan on the stove
48
pick up the ketchup and put it in the basket
Appendix
Table A2: The 10 LIBERO-90 tasks used for the flow-matching evaluation. Each task is listed with its LIBERO-90 index and corresponding language instruction.
Flow Matching
π0.5
Fast-WAM
Source
trained in this work
released checkpoint
released checkpoint
Observation history
8 frames
1 frame
1 frame
Steps executed per replan
12
8
10
Evaluation tasks
10 LIBERO-90 and 4 real
4 LIBERO suites, 10 tasks each
4 LIBERO suites, 10 tasks each
Demonstrations per task
50
–
–
Appendix
Table A3: Configurations of the three frozen policies. Execution horizons are reported for the original 1× setting. Demonstration counts are provided only for the Flow Matching policies trained in this work; dashes indicate counts not listed for the released checkpoints.
ρ
0.2
0.3
0.5
1
1.5
2
3
4
fs
0.20
0.30
0.50
1.00
1.35
1.80
2.70
3.60
fc
0.35
0.50
0.80
1.00
1.00
1.00
1.00
1.00
Appendix
Table A4: Scaling endpoints. One pair per commanded rate, shared by every policy and both simulation and the real robot.
Type
Task
1×
1.5×
2×
0.5×
0.3×
Ref.
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
Structure constrained
2
121
100
98
105
119
93
100
242
194
440
284
32
204
148
160
178
127
125
152
364
327
580
489
40
203
158
176
161
161
149
175
360
338
584
560
86
157
127
116
123
109
105
110
279
255
448
409
Bounded receptacle
18
188
138
136
153
125
112
140
336
328
557
497
Appendix
Table A5: Completion steps under fixed-rate execution on 10 LIBERO-90 tasks with Flow Matching. Each entry reports the mean number of completion steps over that condition’s successful rollouts. Tasks are identified by their LIBERO-90 indices and described in Table A2 . Vanilla denotes vanilla resampling. Success rates and realized rates are reported in Table 2 .
(a) Robomimic: speedup
1×
2×
3×
4×
Statistic
Ref.
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
Overall summary: three tasks
Success (%)
74.0
64.0
66.0
76.7
32.0
35.3
60.7
10.0
25.3
55.3
Rate ( × )
1.000
1.811
1.702
1.315
1.755
1.545
1.428
2.124
2.041
1.473
Appendix
Table A6: Fixed-stride execution on Robomimic with frozen Diffusion Policies. Success is aggregated over Lift, Can, and Square using the original behavior-cloning checkpoints, with 50 trials per task and condition. Rate is the equal-task mean of reference/method step ratios on jointly successful trials. Bold marks the highest success rate under each command.
Table A7: Fixed-stride execution on CALVIN with frozen π0.5 (ABC → D). Each stride setting is evaluated over 1000 trials with randomly sampled seeds. Success requires completing all five instructions in a sequence. Realized rates are computed from paired reference/method step ratios, with equal weighting across the 34 task types. Bold marks the highest success rate under each command.
Endpoint
Policy
Task
Command
Realized rate
Ref. success (%)
FreeSpeed success (%)
Minimum
Flow Matching
LIBERO-90, task 48
0.3×
0.35×
92
100
π0.5
LIBERO-10, task 6
0.2×
0.22×
92
92
Fast-WAM
LIBERO-Spatial, task 6
0.2×
0.22×
100
100
Maximum
Flow Matching
LIBERO-90, task 18
2×
1.37×
98
98
π0.5
LIBERO-Goal, task 5
3×
2.17×
100
100
Fast-WAM
LIBERO-Goal, task 0
4×
2.53×
100
100
Appendix
Table A8: Endpoints of the per-task success-preserving rate range. For each policy, we report the lowest and highest realized rates among task-command settings where FreeSpeed’s observed success rate matches or exceeds the same task’s 1× reference. Success rates % are based on 50 rollouts per task and condition. Bold indicates the overall endpoints, 0.22× and 2.53× .
Ref.
constant factor ft
Task
1×
1.0
1.1
1.2
1.3
1.4
1.5
1.6
1.7
1.8
1.9
2.0
Free [-2pt]Speed
34
Succ.
98
100
100
100
100
96
92
84
70
62
30
40
100
Rate
1.000
1.19
1.26
1.32
1.35
1.38
1.40
1.13
0.98
1.06
0.99
1.18
1.36
48
Succ.
92
100
100
94
92
92
96
88
88
88
72
46
98
Rate
1.000
1.05
1.17
1.27
1.31
1.33
1.37
1.45
1.48
1.54
1.45
1.50
1.35
86
Succ.
92
92
94
92
90
88
78
56
56
54
40
40
88
Appendix
Table A9: Constant-factor versus adaptive rescaling. Results on LIBERO-90 tasks 34, 48, and 86 at the 2× command, with 50 rollouts per task and condition. The factor f remains constant throughout each rollout and is swept from 1.0 to 2.0 in increments of 0.1 . The setting f=2 corresponds to w/o CR. FreeSpeed adapts the rescaling factor at each replan using Equation ( 5 ).
Figure A2: Trade-off between success rate and realized rate under constant-factor rescaling. Results are averaged equally across tasks 34, 48, and 86 at the 2× command. The left and right axes show success rate and realized rate, respectively. Horizontal dashed lines indicate the results of FreeSpeed and the 1× reference.
Figure A3: Success-weighted realized rate for each task. The joint score J=Sr^ combines success rate S∈[0,1] and realized rate r^ . Each panel shows one task at the 2× command, with 50 rollouts per condition. The open marker identifies the highest-scoring constant factor within the sweep: 1.3 , 1.8 , and 1.4 for tasks 34, 48, and 86, respectively.
Figure A4: Fixed-stride evaluations shown as curves. Results from Tables 2 and 3 , plotted against the requested execution rate. The upper row shows mean success rates, and the lower row shows realized execution rates. The curves illustrate how each method’s success rate deviates from the 1× reference as the requested rate changes.
1×
2×
3×
4×
0.5×
0.3×
0.2×
Ref.
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
vanilla
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
w/o [-2pt]CR
Free [-2pt]Speed
Task 1: Peach
Succ.
100
70
90
100
50
55
100
45
0
95
70
100
65
100
20
95
Steps
324
219
233
234
182
211
204
206
–
202
545
433
864
665
1252
806
Rate
1.00
1.48
1.39
1.39
1.77
1.53
1.59
1.57
–
1.60
0.59
0.75
0.38
0.49
0.26
0.40
Task 2: Tennis
Appendix
Table A10: Full fixed-rate results on four real-robot tasks with Flow Matching. Entries report success rates in %, mean completion steps over each condition’s successful rollouts, and realized rates relative to the 1× reference. A dash indicates that no trial succeeded. Success averages include all four tasks; completion-step and rate averages include only tasks with successful trials. Bold indicates the highest success rate under each command. Vanilla denotes vanilla resampling, which coincides with w/o CR under deceleration.
Figure A5: Observation distribution shift under the 0.3× command. Panels (a) and (b) repeat the principal-component projection and empirical distance CDF of Figure 6 for deceleration. Distances are averaged over the k=5 nearest training vectors in the full conditioning space. Under deceleration, w/o CR coincides with vanilla resampling, so only three execution rules are shown.
Figure A7: Accelerated rollouts of the peach-placing task. One rollout is shown per condition. The reference is sampled every 4.7 seconds of wall-clock time. Each accelerated row shows four frames under the 2× command and three under the 4× command, sampled across the episode. At 4× , w/o CR releases the peach beside the plate, while vanilla resampling throws it; the final frame of the latter rollout shows the peach in the air.
Figure A8: Decelerated rollouts of the peach-placing task. One rollout is shown per condition, sampled every 5.5 seconds of wall-clock time. The displayed slowdown commands are 0.3× and 0.2× . Under deceleration, vanilla resampling and w/o CR are equivalent, so only w/o CR is shown alongside FreeSpeed.
Figure A9: Accelerated rollouts of the tennis-ball task. One rollout is shown per condition. The reference is sampled every 8.4 seconds of wall-clock time. Each accelerated row shows four frames under the 2× command and three under the 4× command, sampled across the episode. At 2× , vanilla resampling completes the task but nearly topples the can. At 4× , w/o CR drops the ball beside the can, while vanilla resampling leaves it behind the can, where it appears to overlap the opening from this camera angle.
Figure A10: Decelerated rollouts of the tennis-ball task. One rollout is shown per condition, sampled every 5.9 seconds of wall-clock time, with frames after the scene settles omitted. At 0.5× , vanilla resampling carries the ball over the can but releases it onto the table beside it.
Figure A11: Accelerated rollouts of the cube-stacking task. One rollout per condition. The reference is sampled every 4.4 seconds of wall-clock time, and each accelerated row is sampled at four points under the 2× command and three under the 4× one, spread over the episode.
Figure A12: Decelerated rollouts of the cube-stacking task. One rollout per condition, sampled every 8.0 seconds of wall-clock time.
Figure A13: Accelerated rollouts of the almond-pouring task. One rollout per condition. The reference is sampled every 6.0 seconds of wall-clock time, and each accelerated row is sampled at four points under the 2× command and three under the 4× one, spread over the episode.
Figure A14: Decelerated rollouts of the almond-pouring task. One rollout per condition, sampled every 17.2 seconds of wall-clock time.
College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China · School of Computing, Shanghai Jiao Tong University, Shanghai, China