When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models
Organizations: KAIST · GIST · POSTECH · NVIDIA
Abstract
Vision-Language-Action (VLA) models serve as unified policies for robotic manipulation, yet their expensive inference forces robots to pause between policy calls, resulting in stop-and-go execution that interrupts smooth motion and prolongs task completion. Extending the action chunk reduces policy calls and hence these pauses, but predicting farther into the future makes long-chunk execution unreliable. To understand where this unreliability arises, we analyze action errors within long chunks and find that they concentrate around transitions between manipulation subskills, growing sharply with chunk length. This suggests the importance of transition timing, i.e., when to switch subskills within a chunk. Motivated by this observation, we introduce RACE (Reliable Action-Chunk Extension), a framework that predicts the transition timing from an auxiliary one-step denoising pass and conditions action generation on it. By learning and conditioning on transition timing, RACE reduces errors at subskill transitions and enables reliable execution of longer action chunks. Across simulation benchmarks, RACE outperforms fine-tuning at the same chunk length; with 2x longer chunks, it surpasses recent state-of-the-art and efficient VLAs in success rate, and with 4x longer chunks, it remains competitive. On a real robot, RACE uses 4x longer chunks, which reduces the idle time caused by stop-and-go execution by about 5x, while achieving a higher success rate than fine-tuning with the same chunk length. Code and a real-robot demo are available at https://github.com/Seonghoon-Yu/RACE-VLA
Figures & tables
| Method | Spd | In-dist. | Category | Common. | Instruct. | Texture | Avg. | |||||||
| SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | |||
| Baseline | 5 | 40.4 | 56.7 | 21.4 | 35.6 | 17.0 | 33.7 | 18.0 | 35.9 | 26.0 | 42.1 | 24.6 | 40.8 | |
| [0.2pt/1pt] RACE (ours) | 5 | 51.6 | 68.2 | 25.0 | 38.1 | 26.0 | 40.0 | 20.8 | 37.6 | 29.0 | 45.3 | 30.5 | 45.8 | |
| Action-Chunk Extension | ||||||||||||||
| [0.2pt/1pt] Fine-tuning | 10 | 53.0 | 67.8 | 30.2 | 41.8 | 24.0 | 40.5 | 22.0 | 40.4 | 31.8 | 49.0 | 32.2 | 47.9 | |
| 15 | 51.4 | 65.6 | 26.6 | 39.6 | 23.4 | 40.3 | 19.0 | 37.1 | 29.4 | 46.3 | 30.0 | 45.8 | ||
| Method | Spd | PnP | Doors | Drawer | Sink | Stove | Coffee | Micro. | Nav. | Avg. | |
| Baseline | 5 | 51.5 | 57.0 | 95.0 | 74.7 | 40.0 | 65.3 | 79.0 | 22.0 | 60.4 | |
| [0.2pt/1pt] RACE (ours) | 5 | 57.2 | 56.0 | 93.0 | 77.3 | 42.0 | 68.0 | 80.0 | 40.0 | 63.5 | |
| Action-Chunk Extension | |||||||||||
| [0.2pt/1pt] Fine-tuning | 10 | 1.96 | 62.2 | 56.0 | 94.0 | 75.3 | 45.0 | 65.3 | 89.0 | 26.0 | 65.0 |
| 15 | 2.94 | 60.5 | 58.0 | 94.0 | 68.7 | 35.0 | 72.0 | 91.0 | 26.0 | 64.2 | |
| 20 | 3.84 | 56.8 | 61.0 | 89.0 | 62.7 | 38.0 | 62.7 | 87.0 | 18.0 | 60.8 | |
| Transition-timing Prediction Head (Sec. 3.2 ) | Transition-conditioned Action Generation (Sec. 3.3 ) | Training Objectives | Avg. SR | ||
| - | - | - | - | 28.1 | |
| [0.2pt/1pt] - | - | - | 28.4 | ||
| - | - | 30.6 | |||
| - | 32.4 | ||||
| [0.2pt/1pt] | 33.0 | ||||
| Timing prior at inference | Prior shift | ||||
| Random | Zero | Predicted (ours) | |||
| Avg. | 23.5 | 31.8 | 33.0 | 30.2 | 30.7 |
| Method | SR | Time | Idle | |
| Fine-tuning | 5 | 36% | 12.6 s | 2.11 s |
| 20 | 48% | 11.6 s | 0.40 s | |
| [0.2pt/1pt] RACE (ours) | 20 | 66% | 11.8 s | 0.41 s |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | # Episodes | Total steps per episode | Transitions per episode |
| add_condiment | 500 | ||
| insert_flower | 500 | ||
| select_toy | 500 | ||
| select_chemistry_tube | 500 | ||
| select_drink | 500 | ||
| select_poker | 500 |
| Transition-label sources | Transitions per episode | Precision | Recall |
| Random | 5.0 | 0.30 | 0.28 |
| Speed minima ( Nie et al., 2026 ) | 2.4 | 0.40 | 0.21 |
| [0.2pt/1pt] PELT (used) | 3.8 | 0.82 | 0.70 |
| Method | Spd | Spatial | Object | Goal | Long | Avg. | |
| Baseline | 5 | 98.83 | 98.17 | 97.00 | 93.83 | 96.96 | |
| Action-Chunk Extension | |||||||
| [0.2pt/1pt] Fine-tuning | 10 | 97.60 | 98.80 | 97.60 | 94.40 | 97.10 | |
| 15 | 97.40 | 99.00 | 97.80 | 93.60 | 96.95 | ||
| 20 | 94.80 | 97.60 | 94.00 | 91.00 | 94.35 | ||
| [0.2pt/1pt] PolicyTrim ∗ ( Wang et al., 2026c ) | 13.75 | 97.80 | 98.50 | 98.80 | 93.30 | 97.10 | |
| Method | Spd | In-dist. | Category | Common. | Instruct. | Texture | Avg. | ||||||||
| SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | ||||
| Baseline | 5 | 40.4 | 56.7 | 21.4 | 35.6 | 17.0 | 33.7 | 18.0 | 35.9 | 26.0 | 42.1 | 24.6 | 40.8 | ||
| [0.2pt/1pt] RACE (ours) | 5 | 51.6 | 68.2 | 25.0 | 38.1 | 26.0 | 40.0 | 20.8 | 37.6 | 29.0 | 45.3 | 30.5 | 45.8 | ||
| 10 | 57.6 | 70.9 | 30.4 | 41.8 | 23.2 | 39.0 | 25.0 | 42.0 | 36.0 | 52.8 | 34.4 | 49.3 | |||
| 20 | 54.0 | 68.2 | 27.2 | 38.8 | 28.2 | 41.4 | 21.0 | 39.4 | 34.8 | 51.0 | 33.0 | 47.8 | |||
| 30 | 50.2 | 65.9 | 24.4 | 35.9 | 20.0 | 36.1 | 16.4 | 36.4 | 31.4 | 47.0 | 28.5 | 44.3 | |||
| Setting | In-dist. | Category | Common. | Instruct. | Texture | Avg. | ||||||
| SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | |
| VLM frozen (used) | 54.0 | 68.2 | 27.2 | 38.8 | 28.2 | 41.4 | 21.0 | 39.4 | 34.8 | 51.0 | 33.0 | 47.8 |
| VLM joint training | 57.8 | 72.1 | 30.6 | 43.2 | 28.6 | 43.9 | 31.0 | 50.4 | 46.6 | 62.6 | 38.9 | 54.4 |
| Setting | Training step | Learnable params (M) | Time per update (s) | Wall-clock time (h) | Peak memory (GB/GPU) |
| VLM frozen (used) | 40K | 811 | 3.2 | 36 | 18 |
| [0.2pt/1pt] VLM joint training | 40K | 3,734 | 8.6 | 96 | 40 |
| Method | In-dist. | Category | Common. | Instruct. | Texture | Avg. | ||||||||
| SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | |||
| Baseline | 10 | 5 | 40.4 | 56.7 | 21.4 | 35.6 | 17.0 | 33.7 | 18.0 | 35.9 | 26.0 | 42.1 | 24.6 | 40.8 |
| Fine-tuning | 20 | 20 | 48.8 | 65.3 | 22.8 | 37.2 | 20.2 | 38.2 | 20.8 | 39.3 | 27.8 | 43.9 | 28.1 | 44.8 |
| 40 | 20 | 52.6 | 68.2 | 24.0 | 36.5 | 22.4 | 40.5 | 20.8 | 40.9 | 36.0 | 53.7 | 31.2 | 48.0 | |
| [0.2pt/1pt] RACE (ours) | 20 | 20 | 54.0 | 68.2 | 27.2 | 38.8 | 28.2 | 41.4 | 21.0 | 39.4 | 34.8 | 51.0 | 33.0 | 47.8 |
| 40 | 20 | 56.8 | 70.2 | 28.0 | 38.6 | 27.4 | 43.3 | 28.2 | 44.9 | 35.8 | 51.9 | 35.2 | 49.8 | |
| Component | ACoT-VLA | RACE (ours) | |
| VLM | 53.5 ms | 53.5 ms | 53.5 ms |
| Action expert ( i.e. , 10 denoising steps) | 30.8 ms | 30.8 ms | 30.8 ms |
| [0.2pt/1pt] Explicit action reasoner † | – | 32.5 ms | – |
| Implicit action reasoner † | – | 3.8 ms | – |
| [0.2pt/1pt] Aux. one-step denoising pass (Sec. 3.2 ) | – | – | 3.0 ms |
| Transition-timing head (Sec. 3.2 ) | – | – | 0.3 ms |
| Method | Training step | Global batch | Learnable params (M) | Time per update (s) | Wall-clock time (h) | Peak memory (GB/GPU) |
| Fine-tuning | 40K | 64 | 693 | 2.5 | 28 | 17 |
| [0.2pt/1pt] ACoT-VLA † ( Zhong et al., 2026 ) | 60K | 128 | 1,309 | 50.6 | 843 | 33 |
| [0.2pt/1pt] RACE (ours) | 40K | 64 | 811 | 3.2 | 36 | 18 |
| Method | Training step | Global batch | Learnable params (M) | Time per update (s) | Wall-clock time (h) | Peak memory (GB/GPU) |
| Fine-tuning | 20K | 64 | 693 | 2.5 | 14 | 17 |
| [0.2pt/1pt] PolicyTrim † ( Wang et al., 2026c ) | 500 + 500 | 2,048 | 693 | 1,644 | 457 ( 4 = 1,828) | 36 |
| ACoT-VLA ‡ ( Zhong et al., 2026 ) | 20K | 128 | 1,309 | 49.8 | 277 | 33 |
| [0.2pt/1pt] RACE (ours) | 20K | 64 | 811 | 3.0 | 17 | 18 |
| RACE | In-dist. | Category | Common. | Instruct. | Texture | Avg. | |||||||
| SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | ||
| Run 1 (reported) | 20 | 54.0 | 68.2 | 27.2 | 38.8 | 28.2 | 41.4 | 21.0 | 39.4 | 34.8 | 51.0 | 33.0 | 47.8 |
| Run 2 | 20 | 53.6 | 68.9 | 28.8 | 40.7 | 29.0 | 45.0 | 22.8 | 40.4 | 35.0 | 52.3 | 33.8 | 49.5 |
| Run 3 | 20 | 52.4 | 66.8 | 32.2 | 42.8 | 25.8 | 40.9 | 23.8 | 43.1 | 30.4 | 49.3 | 32.9 | 48.6 |
| [0.2pt/1pt] Mean std | 20 | 53.3 0.8 | 68.0 1.1 | 29.4 2.6 | 40.8 2.0 | 27.7 1.7 | 42.4 2.2 | 22.5 1.4 | 41.0 1.9 | 33.4 2.6 | 50.9 1.5 | 33.3 0.5 | 48.6 0.9 |
| Penalty | # of transitions per episode | Avg. SR |
| 2 | 5.10 | 33.1 |
| 8 | 1.82 | 32.8 |
| [0.2pt/1pt] 4 (used) | 3.57 | 33.0 |
| Gate of Eq. ( 1 ) in Sec. 3.1 | Avg. SR |
| None ( ) | 30.7 |
| [0.2pt/1pt] Learnable per step (ours) | 33.0 |
| Transition label | Avg. SR |
| Hard labels | 32.3 |
| [0.2pt/1pt] Soft labels (used) | 33.0 |
| Training strategy | Avg. SR |
| w/o teacher forcing | 32.9 |
| w/o jittering | 32.8 |
| [0.2pt/1pt] Both (ours) | 33.0 |
| Evaluation type | Metric | Training demos. | Held-out demos. | ||||
| 10 | 15 | 20 | 10 | 15 | 20 | ||
| Transition-chunk Detection | AUROC | 96.0 | 95.4 | 94.4 | 89.2 | 89.7 | 86.5 |
| Acc. | 87.1 | 87.0 | 86.0 | 79.5 | 79.6 | 77.0 | |
| [0.2pt/1pt] Transition-point Localization | Acc. within step | 80.3 | 75.4 | 72.1 | 68.2 | 58.5 | 50.6 |
| Acc. within steps | 90.6 | 85.8 | 83.3 | 81.0 | 69.8 | 65.1 | |
| Model | Shift of the transition-timing prior | ||||||||||
| w/o jittering | 28.6 4.2 | 29.8 3.0 | 24.6 8.2 | 24.8 8.0 | 30.1 2.7 | 32.8 | 30.8 2.0 | 26.4 6.4 | 30.4 2.4 | 28.7 4.1 | 31.1 1.7 |
| [0.2pt/1pt] RACE (ours) | 32.1 0.9 | 31.0 2.0 | 29.4 3.6 | 30.3 2.7 | 30.2 2.8 | 33.0 | 30.7 2.3 | 31.2 1.8 | 30.9 2.1 | 30.9 2.1 | 32.1 0.9 |
| Prior at inference | Training | Held-out | ||||||
| Switch | Exact | step | Missed | Switch | Exact | step | Missed | |
| offset | offset | |||||||
| Shift | 47.3 | 85.6 | 2.5 | 55.0 | 82.2 | 7.5 | ||
| Shift | 57.9 | 91.3 | 1.1 | 60.1 | 89.1 | 5.1 | ||
| Shift | 71.9 | 94.5 | 1.3 | 58.9 | 90.0 | 5.6 | ||
| [0.2pt/1pt] Predicted (ours) | 75.0 | 97.3 | 1.0 | 53.8 | 90.8 | 4.1 | ||
| Error component | MSE | Error reduction (%) over Fine-tuning | |
| Fine-tuning | RACE (ours) | ||
| Translation | 3.448 | 2.995 | 13.1 |
| Rotation | 0.277 | 0.190 | 31.5 |
| Gripper | 0.124 | 0.094 | 24.6 |
| [0.2pt/1pt] Overall (7-dim) | 1.614 | 1.378 | 14.6 |
| Demos. | Method | Exact | step | steps | Missed | Mean error (steps) | False switch |
| Training | Fine-tuning | 70.1 | 94.1 | 97.4 | 2.5 | 0.32 | 0.5 |
| RACE (ours) | 75.2 | 97.2 | 99.0 | 0.9 | 0.26 | 0.6 | |
| Difference | [3.0, 6.9] | [2.2, 4.0] | [1.0, 2.2] | [ , ] | [ , ] | [ , 0.2] | |
| Held-out | Fine-tuning | 50.6 | 86.9 | 93.7 | 6.3 | 0.54 | 0.6 |
| RACE (ours) | 53.8 | 91.5 | 95.9 | 3.9 | 0.50 | 0.7 | |
| Difference | [ , 7.6] | [1.9, 7.1] | [0.5, 4.1] | [ , ] | [ , 0.01] | [ , 0.3] |
| RACE (ours) | |||
| # of policy calls per episode | 45 | 17 | 17 |
| Actions executed per call | 7.0 | 19.9 | 20.1 |
| [0.2pt/1pt] Latency per call | 271 ms | 285 ms | 372 ms |
| Camera capture | 122 ms | 112 ms | 109 ms |
| Network | 49 ms | 64 ms | 75 ms |
| Model | Env. steps per episode | Policy calls per episode | Spd | ||
| Pooled | Task-balanced | Range across tasks | |||
| Baseline | 117.9 | 23.93 | – | ||
| [0.2pt/1pt] Fine-tuning | 112.3 | 6.10 | – | ||
| RACE (ours) | 113.6 | 6.16 | – | ||
| Episode set | # Episodes | Fine-tuning | RACE (ours) |
| Common successes with the baseline (used) | 228 / 256 | ||
| [0.2pt/1pt] Common successes of all three models | 127 | ||
| with each task weighted equally | 127 | ||
| All episodes, regardless of success | 2,500 |
| RACE (ours) | ||
| Error ratio (transition / matched) | 1.34 [1.15, 1.56] | 1.30 [1.05, 1.62] |