Vision-Language-Action (VLA) models serve as unified policies for robotic manipulation, yet their expensive inference forces robots to pause between policy calls, resulting in stop-and-go execution that interrupts smooth motion and prolongs task completion. Extending the action chunk reduces policy calls and hence these pauses, but predicting farther into the future makes long-chunk execution unreliable. To understand where this unreliability arises, we analyze action errors within long chunks and find that they concentrate around transitions between manipulation subskills, growing sharply with chunk length. This suggests the importance of transition timing, i.e., when to switch subskills within a chunk. Motivated by this observation, we introduce RACE (Reliable Action-Chunk Extension), a framework that predicts the transition timing from an auxiliary one-step denoising pass and conditions action generation on it. By learning and conditioning on transition timing, RACE reduces errors at subskill transitions and enables reliable execution of longer action chunks. Across simulation benchmarks, RACE outperforms fine-tuning at the same chunk length; with 2x longer chunks, it surpasses recent state-of-the-art and efficient VLAs in success rate, and with 4x longer chunks, it remains competitive. On a real robot, RACE uses 4x longer chunks, which reduces the idle time caused by stop-and-go execution by about 5x, while achieving a higher success rate than fine-tuning with the same chunk length. Code and a real-robot demo are available at https://github.com/Seonghoon-Yu/RACE-VLA
Figures & tables
Figure 1: Effect of chunk length H on VLABench, with per-episode inference speedup over π0.5 (Appendix E.2 ).
Figure 2: Action prediction errors around subskill transitions: (a) an example of transitions between manipulation subskills; (b) action errors spike at the transition point within a chunk, and these spikes grow with the chunk length H ; and (c) RACE reduces these error spikes across chunk lengths H .
Figure 3: Overview of RACE. Given an observation and instruction, the frozen VLM extracts features cached for both passes. RACE predicts a transition-timing prior via an auxiliary one-step denoising pass (Sec. 3.2 ), then restarts denoising from the same noise, conditioned on it (Sec. 3.3 ).
Method
Hexec
Spd ↑
In-dist.
Category
Common.
Instruct.
Texture
Avg.
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
π0.5 Baseline
5
1.00×
40.4
56.7
21.4
35.6
17.0
33.7
18.0
35.9
26.0
42.1
24.6
40.8
[0.2pt/1pt] RACE (ours)
5
0.96×
51.6
68.2
25.0
38.1
26.0
40.0
20.8
37.6
29.0
45.3
30.5
45.8
Action-Chunk Extension
[0.2pt/1pt] π0.5 Fine-tuning
10
2.01×
53.0
67.8
30.2
41.8
24.0
40.5
22.0
40.4
31.8
49.0
32.2
47.9
15
2.83×
51.4
65.6
26.6
39.6
23.4
40.3
19.0
37.1
29.4
46.3
30.0
45.8
Table 1: Comparison on VLABench. We report the success rate (SR) and the progress score (PS), the fraction of completed sub-tasks. Spd is the per-episode inference speedup relative to the π0.5 baseline (Appendix E.2 ). Hexec is the number of actions executed per policy call, and † denotes a model fine-tuned to execute that many actions. ‡ indicates the average execution length, since adaptive chunking methods truncate each chunk adaptively. Results for a wider range of chunk lengths are in Appendix B.2 .
Method
Hexec
Spd ↑
PnP
Doors
Drawer
Sink
Stove
Coffee
Micro.
Nav.
Avg.
π0.5 Baseline
5
1.00×
51.5
57.0
95.0
74.7
40.0
65.3
79.0
22.0
60.4
[0.2pt/1pt] RACE (ours)
5
0.97×
57.2
56.0
93.0
77.3
42.0
68.0
80.0
40.0
63.5
Action-Chunk Extension
[0.2pt/1pt] π0.5 Fine-tuning
10
1.96 ×
62.2
56.0
94.0
75.3
45.0
65.3
89.0
26.0
65.0
15
2.94 ×
60.5
58.0
94.0
68.7
35.0
72.0
91.0
26.0
64.2
20
3.84 ×
56.8
61.0
89.0
62.7
38.0
62.7
87.0
18.0
60.8
Table 2: Comparison on RoboCasa-H50. We report the success rate on eight task categories, with the average weighted by the number of tasks per category. Spd is the per-episode inference speedup relative to the π0.5 baseline. ‡ indicates the average execution length, since adaptive chunking methods truncate each chunk adaptively.
Table 3: Ablation of each component and training objective. The auxiliary one-step pass is always performed when the transition-timing prediction head is present, since the head reads its features; Laux only controls whether this pass is additionally supervised as a denoiser.
Table 4: Ablation within the transition-timing prediction head (Sec. 3.2 ). Analyses of transition-label validity and sensitivity to PELT settings are provided in Appendices A.5 and C.1 , respectively.
Figure 4: Errors at transition and non-transition actions with Hexec=10 and 20 on VLABench, where π0.5 is fine-tuned to execute Hexec actions per policy call. At each position within a chunk, we compare the error of actions at transition points (transition error) with that of the remaining actions (non-transition error), which we use as a reference for the error growth with prediction horizon.
Figure 6: Adjacent action-token similarity.
Timing prior at inference
Prior shift
Random
Zero
Predicted (ours)
−1
+1
Avg.
23.5
31.8
33.0
30.2
30.7
Table 5: Replacement and shift of timing prior.
Method
Hexec
SR ↑
Time ↓
Idle ↓
π0.5 Fine-tuning
5
36%
12.6 s
2.11 s
20
48%
11.6 s
0.40 s
[0.2pt/1pt] RACE (ours)
20
66%
11.8 s
0.41 s
Table 6: Results on a real robot.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Visualization of subskill transitions. Change points detected by PELT ( Killick et al., 2012 ) are overlaid on the action sequences of two VLABench demonstrations, alongside a representative frame for each subskill. Subskill names are generated with Claude Code ( Anthropic, 2025 ) for illustration only; only the transition points are used as labels for policy learning, not the semantics.
Task
# Episodes
Total steps per episode
Transitions per episode
add_condiment
500
167.6±14.1
5.59±0.80
insert_flower
500
167.3±15.6
5.42±0.87
select_toy
500
160.2±13.7
4.02±1.09
select_chemistry_tube
500
78.4±8.6
3.68±0.69
select_drink
500
97.4±8.0
3.31±0.60
select_poker
500
75.4±7.0
3.06±0.47
Appendix
Table 7: Statistics of detected subskill transitions on VLABench training demonstrations. For each task, we report the number of steps and the number of subskill transitions ( i.e. , change points) per episode as mean ± standard deviation.
Figure 8: An example of transition-timing target and prediction on a held-out VLABench demonstration ( Hexec=20 ). Top: standardized action signal and PELT change points; the pink-shaded chunk is enlarged in the bottom plot. Bottom: soft label pt and predicted prior p^t in this chunk.
Figure 9: Qualitative comparison between PELT change points and reference subskill boundaries recorded by the VLABench demonstration generator.
Transition-label sources
Transitions per episode
Precision
Recall
Random
5.0
0.30
0.28
Speed minima ( Nie et al., 2026 )
2.4
0.40
0.21
[0.2pt/1pt] PELT (used)
3.8
0.82
0.70
Appendix
Table 8: Agreement between transition labels and the reference subskill boundaries recorded by the VLABench demonstration generator, on 300 newly generated demonstrations with a matching tolerance of ±3 steps. Precision and recall are counted per transition and pooled over all episodes.
Method
Hexec
Spd ↑
Spatial
Object
Goal
Long
Avg.
π0.5 Baseline
5
1.00×
98.83
98.17
97.00
93.83
96.96
Action-Chunk Extension
[0.2pt/1pt] π0.5 Fine-tuning
10
1.98×
97.60
98.80
97.60
94.40
97.10
15
2.96×
97.40
99.00
97.80
93.60
96.95
20
3.91×
94.80
97.60
94.00
91.00
94.35
[0.2pt/1pt] PolicyTrim ∗ ( Wang et al., 2026c )
13.75
4.85×
97.80
98.50
98.80
93.30
97.10
Appendix
Table 9: Comparison on LIBERO. ∗ PolicyTrim trains a separate model for each suite, and its Hexec is averaged over the four suites. Its training cost is also substantially higher; see Appendix B.7 . Spd is the per-episode inference speedup relative to the π0.5 baseline. ‡ indicates the average execution length, since adaptive chunking methods truncate each chunk adaptively.
Method
Hexec
Spd ↑
In-dist.
Category
Common.
Instruct.
Texture
Avg.
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
π0.5 Baseline
5
1.00×
40.4
56.7
21.4
35.6
17.0
33.7
18.0
35.9
26.0
42.1
24.6
40.8
[0.2pt/1pt] RACE (ours)
5
0.96×
51.6
68.2
25.0
38.1
26.0
40.0
20.8
37.6
29.0
45.3
30.5
45.8
10
1.92×
57.6
70.9
30.4
41.8
23.2
39.0
25.0
42.0
36.0
52.8
34.4
49.3
20
3.68×
54.0
68.2
27.2
38.8
28.2
41.4
21.0
39.4
34.8
51.0
33.0
47.8
30
5.11×
50.2
65.9
24.4
35.9
20.0
36.1
16.4
36.4
31.4
47.0
28.5
44.3
Appendix
Table 10: Results across a wide range of execution chunk lengths on VLABench.
Setting
In-dist.
Category
Common.
Instruct.
Texture
Avg.
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
VLM frozen (used)
54.0
68.2
27.2
38.8
28.2
41.4
21.0
39.4
34.8
51.0
33.0
47.8
VLM joint training
57.8
72.1
30.6
43.2
28.6
43.9
31.0
50.4
46.6
62.6
38.9
54.4
Appendix
Table 11: Effect of training the VLM jointly on VLABench with Hexec=20 . VLM frozen indicates our post-training setting, where the VLM is kept frozen and only the action expert is post-trained. VLM joint training means that the VLM is also updated.
Setting
Training step
Learnable params (M)
Time per update (s)
Wall-clock time (h)
Peak memory (GB/GPU)
VLM frozen (used)
40K
811
3.2
36
18
[0.2pt/1pt] VLM joint training
40K
3,734
8.6
96
40
Appendix
Table 12: Training cost of the VLM-frozen and VLM joint training settings on VLABench with Hexec=20 . Both models are trained for 40K steps on four RTX A6000 GPUs. Peak memory is the maximum allocated memory per GPU.
Method
H
Hexec
In-dist.
Category
Common.
Instruct.
Texture
Avg.
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
π0.5 Baseline
10
5
40.4
56.7
21.4
35.6
17.0
33.7
18.0
35.9
26.0
42.1
24.6
40.8
π0.5 Fine-tuning
20
20
48.8
65.3
22.8
37.2
20.2
38.2
20.8
39.3
27.8
43.9
28.1
44.8
40
20
52.6
68.2
24.0
36.5
22.4
40.5
20.8
40.9
36.0
53.7
31.2
48.0
[0.2pt/1pt] RACE (ours)
20
20
54.0
68.2
27.2
38.8
28.2
41.4
21.0
39.4
34.8
51.0
33.0
47.8
40
20
56.8
70.2
28.0
38.6
27.4
43.3
28.2
44.9
35.8
51.9
35.2
49.8
Appendix
Table 13: Results with different prediction chunk lengths H at the same execution chunk length Hexec=20 on VLABench, with the π0.5 baseline as reference. Only the first Hexec of the H predicted actions are executed per policy call.
Table 14: Generalization to other VLA baselines on LIBERO.
Table 15: Per-call inference latency, measured on a single NVIDIA RTX A6000 GPU. † denotes ACoT-VLA components measured in our setup with its official implementation, where the explicit action reasoner uses an additional 10-step denoising process to generate coarse action trajectories.
Method
Training step
Global batch
Learnable params (M)
Time per update (s)
Wall-clock time (h)
Peak memory (GB/GPU)
π0.5 Fine-tuning
40K
64
693
2.5
28
17
[0.2pt/1pt] ACoT-VLA † ( Zhong et al., 2026 )
60K
128
1,309
50.6
843
33
[0.2pt/1pt] RACE (ours)
40K
64
811
3.2
36
18
Appendix
Table 16: Training efficiency comparison on VLABench. All methods are measured on four NVIDIA RTX A6000 GPUs, with the compared methods run from their official implementations. † : following its official setting, ACoT-VLA freezes the LLM backbone and trains the visual encoder, the action expert, and its components.
Method
Training step
Global batch
Learnable params (M)
Time per update (s)
Wall-clock time (h)
Peak memory (GB/GPU)
π0.5 Fine-tuning
20K
64
693
2.5
14
17
[0.2pt/1pt] PolicyTrim † ( Wang et al., 2026c )
500 + 500
2,048
693
1,644
457 ( × 4 = 1,828)
36
ACoT-VLA ‡ ( Zhong et al., 2026 )
20K
128
1,309
49.8
277
33
[0.2pt/1pt] RACE (ours)
20K
64
811
3.0
17
18
Appendix
Table 17: Training efficiency comparison on LIBERO. All methods are measured on four NVIDIA RTX A6000 GPUs, with PolicyTrim and ACoT-VLA run from their official implementations. † : PolicyTrim trains a separate model for each of the four LIBERO suites with two-stage reinforcement learning (500 + 500 RL iterations); its values are for a single suite, and the parentheses give the estimated total for all four suites. Its time per update is per RL iteration, including rollout collection. ‡ : following its official setting, ACoT-VLA freezes the LLM backbone and trains the visual encoder, the action expert, and its components.
RACE
Hexec
In-dist.
Category
Common.
Instruct.
Texture
Avg.
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
SR
PS
Run 1 (reported)
20
54.0
68.2
27.2
38.8
28.2
41.4
21.0
39.4
34.8
51.0
33.0
47.8
Run 2
20
53.6
68.9
28.8
40.7
29.0
45.0
22.8
40.4
35.0
52.3
33.8
49.5
Run 3
20
52.4
66.8
32.2
42.8
25.8
40.9
23.8
43.1
30.4
49.3
32.9
48.6
[0.2pt/1pt] Mean ± std
20
53.3 ± 0.8
68.0 ± 1.1
29.4 ± 2.6
40.8 ± 2.0
27.7 ± 1.7
42.4 ± 2.2
22.5 ± 1.4
41.0 ± 1.9
33.4 ± 2.6
50.9 ± 1.5
33.3 ± 0.5
48.6 ± 0.9
Appendix
Table 18: Statistical significance of RACE on VLABench with Hexec=20 . In addition to the model reported in Tab. 1 (Run 1), we train RACE twice more with different random seeds (Runs 2 and 3). The last row reports the mean and standard deviation over the three runs.
Penalty β
# of transitions per episode
Avg. SR
2
5.10
33.1
8
1.82
32.8
[0.2pt/1pt] 4 (used)
3.57
33.0
Appendix
Table 19: Sensitivity to the PELT penalty on VLABench with Hexec=20 . We vary the penalty β in Eq. ( 8 ) of Appendix A.3 .
Gate αk of Eq. ( 1 ) in Sec. 3.1
Avg. SR
None ( αk=1 )
30.7
[0.2pt/1pt] Learnable per step (ours)
33.0
Appendix
Table 20: Effect of the per-step learnable gate on VLABench with Hexec=20 .
Transition label
Avg. SR
Hard labels
32.3
[0.2pt/1pt] Soft labels (used)
33.0
Appendix
Table 21: Effect of soft transition labels on VLABench with Hexec=20 . Hard labels mark only the detected transition step; soft labels spread each transition over neighboring steps with a Gaussian profile (Appendix A.4 ).
Training strategy
Avg. SR
w/o teacher forcing
32.9
w/o jittering
32.8
[0.2pt/1pt] Both (ours)
33.0
Appendix
Table 22: Effect of teacher forcing and jittering on VLABench with Hexec=20 . The robustness provided by jittering against inference-time transition shifts is evaluated separately in Tab. 24 of Appendix D.2 .
Evaluation type
Metric
Training demos.
Held-out demos.
Hexec
Hexec
10
15
20
10
15
20
Transition-chunk Detection
AUROC
96.0
95.4
94.4
89.2
89.7
86.5
Acc.
87.1
87.0
86.0
79.5
79.6
77.0
[0.2pt/1pt] Transition-point Localization
Acc. within ±1 step
80.3
75.4
72.1
68.2
58.5
50.6
Acc. within ±3 steps
90.6
85.8
83.3
81.0
69.8
65.1
Appendix
Table 23: Accuracy (%) of the transition-timing prediction head on VLABench across execution chunk lengths Hexec . Transition-chunk Detection evaluates, for each chunk, whether it contains a transition. Transition-point Localization evaluates, for each chunk containing a transition, where the transition occurs. Multi-transition Localization evaluates how many of the transitions within a chunk are captured, counting every predicted peak rather than only the highest one. Results are reported on both training and held-out demonstrations.
Model
Shift δ of the transition-timing prior
−15
−10
−5
−3
−1
0
+1
+3
+5
+10
+15
w/o jittering
28.6 − 4.2
29.8 − 3.0
24.6 − 8.2
24.8 − 8.0
30.1 − 2.7
32.8
30.8 − 2.0
26.4 − 6.4
30.4 − 2.4
28.7 − 4.1
31.1 − 1.7
[0.2pt/1pt] RACE (ours)
32.1 − 0.9
31.0 − 2.0
29.4 − 3.6
30.3 − 2.7
30.2 − 2.8
33.0
30.7 − 2.3
31.2 − 1.8
30.9 − 2.1
30.9 − 2.1
32.1 − 0.9
Appendix
Table 24: Sensitivity to shifts of the predicted transition timing at inference on VLABench with Hexec=20 . The peaks of the predicted prior are shifted by δ steps within the chunk; peaks shifted beyond the chunk are dropped.
Prior at inference
Training
Held-out
Switch
Exact ↑
±1 step ↑
Missed ↓
Switch
Exact ↑
±1 step ↑
Missed ↓
offset
offset
Shift k=−3
−0.64
47.3
85.6
2.5
−0.47
55.0
82.2
7.5
Shift k=−2
−0.48
57.9
91.3
1.1
−0.26
60.1
89.1
5.1
Shift k=−1
−0.25
71.9
94.5
1.3
+0.06
58.9
90.0
5.6
[0.2pt/1pt] Predicted (ours)
−0.05
75.0
97.3
1.0
+0.29
53.8
90.8
4.1
Appendix
Table 25: Effect of the transition-timing prior on the generated gripper switches on VLABench with Hexec=20 . The predicted prior is shifted by k steps or replaced with zeros at inference, over chunks where the demonstrated gripper command switches exactly once. Switch offset is the predicted minus the demonstrated switch step. Values under the predicted prior differ slightly from Tab. 27 due to different noise draws.
Error component
MSE
Error reduction (%) over π0.5 Fine-tuning
π0.5 Fine-tuning
RACE (ours)
Translation
3.448
2.995
13.1
Rotation
0.277
0.190
31.5
Gripper
0.124
0.094
24.6
[0.2pt/1pt] Overall (7-dim)
1.614
1.378
14.6
Appendix
Table 26: Action errors (MSE) at transitions by action component on VLABench with Hexec=20 .
Demos.
Method
Exact ↑
±1 step ↑
±3 steps ↑
Missed ↓
Mean error (steps) ↓
False switch ↓
Training
π0.5 Fine-tuning
70.1
94.1
97.4
2.5
0.32
0.5
RACE (ours)
75.2
97.2
99.0
0.9
0.26
0.6
Difference
+5.1 [3.0, 6.9]
+3.1 [2.2, 4.0]
+1.6 [1.0, 2.2]
−1.6 [ −2.2 , −1.0 ]
−0.06 [ −0.08 , −0.04 ]
+0.0 [ −0.1 , 0.2]
Held-out
π0.5 Fine-tuning
50.6
86.9
93.7
6.3
0.54
0.6
RACE (ours)
53.8
91.5
95.9
3.9
0.50
0.7
Difference
+3.2 [ −1.2 , 7.6]
+4.6 [1.9, 7.1]
+2.2 [0.5, 4.1]
−2.4 [ −4.3 , −0.9 ]
−0.04 [ −0.09 , 0.01]
+0.1 [ −0.1 , 0.3]
Appendix
Table 27: Gripper switching timing on VLABench with Hexec=20 , over chunks where the demonstrated gripper command switches exactly once, evaluated on 500 randomly sampled training demonstrations and the 300 held-out demonstrations of Appendix A.5 . All values are in % except Mean error , which is in action steps (0.1 s each). Mean error is the mean absolute offset between the predicted and the demonstrated switch, over chunks in which the model predicts a switch; missed switches are counted in the Missed column. False switch is the fraction of chunks without a demonstrated switch in which the prediction switches. Difference is RACE minus fine-tuning on the same chunks, with 95% percentile confidence intervals of this paired difference in brackets, obtained from 1,000 bootstrap resamples of episodes.
π0.5
RACE (ours)
Hexec=5
Hexec=20
Hexec=20
# of policy calls per episode
45
17
17
Actions executed per call
7.0
19.9
20.1
[0.2pt/1pt] Latency per call
271 ms
285 ms
372 ms
Camera capture
122 ms
112 ms
109 ms
Network
49 ms
64 ms
75 ms
Appendix
Table 28: Breakdown of the idle time in Tab. 6 over successful trials. The robot keeps executing the current chunk while the next one is prepared, and pauses only when the latency exceeds this execution buffer, i.e. , the execution time of the actions remaining in a chunk when it arrives, after real-time chunking skips the actions elapsed during inference. Latencies are averaged over calls and, unlike Tab. 15 for latency overhead comparisons, include camera capture, network transfer, and real-time chunking ( Black et al., 2025 ) . Pause per call exceeds Latency − buffer because calls within the buffer count as zero pause.
Figure 10: Task Visualization. Keyframes of a successful rollout of RACE from the front view.
Model
Env. steps per episode
Policy calls per episode
Spd ↑
Pooled
Task-balanced
Range across tasks
π0.5 Baseline
117.9
23.93
1.00×
1.00×
–
[0.2pt/1pt] π0.5 Fine-tuning
112.3
6.10
3.92×
3.89×
3.15 – 4.70×
RACE (ours)
113.6
6.16
3.74×
3.64×
2.54 – 4.53×
Appendix
Table 29: Breakdown of the inference speedup at Hexec=20 on VLABench, over the 127 episodes that all three models complete successfully. Spd is computed over all episodes pooled and with each task weighted equally; the last column shows its range across tasks.
Episode set
# Episodes
π0.5 Fine-tuning
RACE (ours)
Common successes with the π0.5 baseline (used)
228 / 256
3.82×
3.68×
[0.2pt/1pt] Common successes of all three models
127
3.92×
3.74×
with each task weighted equally
127
3.89×
3.64×
All episodes, regardless of success
2,500
4.00×
3.94×
Appendix
Table 30: Inference speedup at Hexec=20 on VLABench under different sets of episodes. The first row is the setting used in the main comparisons.
π0.5
RACE (ours)
Error ratio (transition / matched)
1.34 [1.15, 1.56]
1.30 [1.05, 1.62]
Appendix
Table 31: Ratio of the action error at transitions to that at non-transition positions matched by position within the chunk and by decile of the second-order action difference ch , on VLABench with Hexec=20 . A ratio of 1 means no extra error at transitions. Brackets denote 95% episode-level bootstrap confidence intervals.
Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.
Yushan Liu, Peibo Sun, Xintao Chao +8
Tsinghua University · Shanghai Jiao Tong University · Peking University +2
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2
Vision-Language-Action (VLA) policies that execute fixed-length action chunks can exhibit multimodal bifurcation: a cross-chunk inconsistency in which adjacent chunks generated from independent Gaussian latents can converge to incompatible trajectory modes, producing abrupt discontinuities at chunk boundaries. Existing remedies either require backpropagation through the policy at each denoising step, rely on rejection sampling, or require retraining, each trading computational cost or task reliability for smoother transitions. We propose SEAM (Smooth Execution of Action-Chunked Motion), a training-free inference-time method for flow matching VLAs. SEAM exploits a simple synchronous-execution insight: after the robot consumes the executed prefix, the previous chunk's unexecuted tail is already available as an analytic consistency reference. Its core mechanism, Velocity-guided Loss Steering (VLS), derives a time-dependent target from this tail and applies a closed-form correction after each Euler step without backpropagating through the policy network. On LIBERO-10 with pi_0.5, SEAM reduces boundary jerk by 28%, reduces chunk transition discontinuity by 27%, preserves baseline-level task success, and keeps denoising-loop cost near the unguided baseline.