While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics. Measuring these failures is challenging because action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes. We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails. To manage this, we introduce CARE, an approach for certified accelerator selection. CARE uses paired rollouts on a calibration set to provide finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It deploys the fastest certified candidate, falling back to the reference if none qualify. By relying only on terminal outcomes and measured compute, CARE applies unchanged across diverse acceleration mechanisms, while sequential testing and failure-triggered reference rollouts keep certification affordable. On four LIBERO suites with OpenVLA-OFT, CARE certifies 9.0--10.8× speedups while guaranteeing (at 95% confidence) that at least 85.8% of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to 75% of trials, whereas CARE stays within budget and its sequential form uses 78.9% fewer rollouts than exhaustive evaluation. CARE further generalizes to flow-step reduction for π0.5, and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.
Figures & tables
Figure 1: Overview of care . Given a reference policy π0 and accelerated candidates Λ={λ1,…,λM} , care first profiles candidate latency on a disjoint profiling set and orders candidates from fastest to slowest. On a separate calibration set, candidates are evaluated sequentially in task-balanced rounds, with each round containing the same number of episodes from every task. For each paired calibration episode, the reference and accelerated policy start from the same evaluation unit U , and care records the episode-level loss Lλ(U)=Y0(U)(1−Yλ(U)) , which equals one only when the reference succeeds and the accelerated policy fails. An anytime-valid test with family-wise error control certifies candidates while allowing early stopping. To reduce calibration cost, the reference policy is executed only when a candidate fails, and its outcome is cached for reuse. care returns the first, and therefore fastest, certified candidate; if none can be certified, it retains π0 . The returned policy satisfies the finite-sample deployment guarantee without adding overhead at inference.
Suite
Selected policy
Speedup
Preserved
Goal
FastV-25 H8
9.02×
93.5%
Spatial
FastV-75 H8
10.08×
95.7%
Object
Plain H8
9.56×
92.2%
Long
VLA-Cache H8
10.81×
85.8%
Table 1: Certified CARE deployments. Speedup is relative to H1; preserved is the simultaneous 95% lower bound on the fraction of H1-solved episodes retained.
Method
α=2.0%
α=2.5%
Latency only
75.0
25.0
Average success
75.0
25.0
Validation tuned
36.8
13.0
Unpaired test
67.1
19.5
care
0.0
0.0
Table 2: Over-budget rate (%) under tighter AIF budgets. care has no observed over-budget deployments, whereas the other selectors exceed the budget in 13.0 – 75.0% of trials.
Method
Mean speedup ( × )
Fallback (%)
Mean rollouts
care , fixed- n
8.88
11.7
2,250
Pareto Testing
6.16
42.2
2,250
Cost-order fixed-sequence LTT
6.35
45.5
2,250
Safe policy improvement
5.30
60.5
2,250
care , sequential
7.32
26.9
592
Table 3: Deployment speed and evaluation cost. Fixed- n care achieves the highest mean speedup; sequential care reduces evaluation cost by 73.7% while keeping a 7.32× speedup and less fallback than all baselines.
Method
Suite
Acceleration
Horizon
Candidate rollouts
Reference rollouts
Total rollouts
Rollout saving (%)
Full-sample HB
Goal
FastV-50
H8
24,000
1,000
25,000
0.0
Spatial
VLA-Cache
H2
Object
VLA-Pruner
H8
Long
Reference
–
Sequential HB
Goal
Reference
–
7,150
156
7,306
70.8
Spatial
Reference
–
Table 4: Evaluation with a broader acceleration candidate family. Reference denotes fallback to H1. care is the only method that certifies an accelerated policy on every suite, using 5,263 rollouts, 78.9% fewer than full-sample HB and fewer than every sequential baseline. Rollouts are aggregated over the four LIBERO suites.
Suite
Selected policy
Compute (% ref.)
Speedup ( × )
Total rollouts
Rollout saving (%)
Goal
Four-step
40
1.26
305
69.5
Spatial
Two-step
20
1.45
254
74.6
Object
Two-step
20
1.45
253
74.7
Long
Reference
100
1.00
464
53.6
Total
1,276
68.1
Table 5: Certified flow-step selection for π0.5 on LIBERO. care certifies reduced-step policies on Goal, Spatial, and Object and retains the ten-step reference on Long, using 1,276 rather than 4,000 rollouts ( 68.1% fewer). Compute and speedup are relative to the ten-step reference.
Model
Task
Certified seeds
Compute (% ref.)
Qwen3.5-9B
Wood
3/3
69.4
Qwen3.5-9B
Drink
2/3
85.8
Llama-3.1-8B
Wood
3/3
74.9
Llama-3.1-8B
Drink
2/3
98.6
Table 6: Certified compute reduction for Crafter agents. collect_wood schedules certify on all seeds at 69.4%–74.9% of reference compute; collect_drink schedules certify on two of three with compute reduced.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Suite
One accelerated decision
Whole-episode AIF
Goal
2.3%
2.6%
Spatial
0.7%
0.8%
Object
4.3%
1.0%
Long
9.5%
2.7%
Appendix
Table 7: Local acceleration effects do not determine whole-episode AIF. The relationship varies across suites because an accelerated action can alter subsequent states, observations, and actions.
Selection rule
Over-budget deployment
Uncorrected empirical selection
46.0%
Family-corrected testing
≤0.15%
Appendix
Table 8: Family correction prevents selection-induced risk inflation. Searching over ten candidates without correction deploys an over-budget policy in 46.0% of trials, whereas family-corrected testing remains at or below 0.15% .
Task-risk profile
Balanced rounds
Task by task
Eight safe, two risky; R=0.050
1.84×10−4
1.00
Graded; R=0.050
4.37×10−4
1.42×10−2
Homogeneous; R=0.050
4.77×10−4
4.77×10−4
One high-risk task; R=0.060
1.08×10−6
1.00
Six safe, four risky; R=0.052
2.43×10−4
1.00
Appendix
Table 9: Balanced acquisition preserves sequential validity. Balanced rounds remain below the candidate-level error budget for every tested task-risk profile, whereas task-by-task ordering can certify an over-budget policy before its high-risk tasks are evaluated.
Procedure
Candidate rollouts
H1 rollouts
Total rollouts
Wall-clock time
Fixed- n test-all
8,000
1,000
9,000
38.6 h
Fixed- n + failure-triggered H1
8,000
349
8,349
33.1 h
Sequential fastest-first + full H1
1,950
1,000
2,950
16.3 h
care
1,950
47
1,997
7.0 h
Appendix
Table 10: Component-wise ablation of certification cost. Sequential fastest-first evaluation reduces candidate rollouts, while failure-triggered evaluation reduces H1 rollouts. Combining both components reduces total rollouts by 77.8% and wall-clock time by 81.9% relative to fixed- n test-all evaluation.
α(%)
n
Over budget (%)
Fallback (%)
Speedup ( × )
2.0
250
0.00
100.0
1.00
2.5
250
0.00
100.0
1.00
3.0
250
0.00
69.5
3.51
5.0
250
0.00
11.7
8.88
10.0
250
0.00
0.0
11.25
2.5
500
0.02
56.3
4.59
Appendix
Table 11: Sensitivity to risk budget and calibration size. A larger risk budget or more calibration evidence reduces fallback and increases speedup, while care maintains an observed over-budget rate of at most 0.02% .
Figure 2: LIBERO task suites. Initial scenes of five of the ten tasks in each suite, rendered from the third-person agent-view camera, with the language instruction below each scene. Spatial instructions have the form “pick up the black bowl … and place it on the plate” and Object instructions “pick up the … and place it in the basket”; only the distinguishing part is shown. Each evaluation unit fixes a task, one initial scene, and the random seeds.
Figure 3: Crafter tasks. Example episodes of collect_wood (top) and collect_drink (bottom). The status bar shows health, food, drink, and energy, followed by the inventory; an episode succeeds once wood enters the inventory or the agent drinks water. Frames are rendered from the environment with a scripted policy for illustration and are not rollouts of the evaluated language-model agents.
Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself. Success rates alone cannot establish whether such gains come from better task execution or from evaluation flaws. We therefore investigate two kinds of benchmark flaws behind these gains: bugs, where the implementation does not match the intended task or evaluation protocol, and design limitations, where success criteria and simulation settings do not fully capture how acceleration affects task execution. Starting from tasks with anomalous gains, we localize root causes by plotting object trajectories against checker acceptance regions, and classify the resulting bugs into task consistency, initialization, and reproducibility. Extending this audit to seven benchmarks, including RoboTwin, LIBERO-Plus, and VLABench, we identify 22 bugs of these types and 4 design limitations. For the latter, we revise permissive success checkers, correct unrealistic object masses, and add a motion-aware score that favors smoother actions. Experiments show that bug fixes can reverse method rankings, moving the baseline from last to first on one task. Addressing design limitations can likewise remove anomalous gains: on another task, the baseline moves from 21 percentage points behind an accelerated method to 5 points ahead. Gains attributed to acceleration can therefore be artifacts of the benchmark rather than better task execution. We release our bug fixes and revised benchmark settings to support trustworthy evaluation of VLA acceleration.
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2
Deploying vision--language--action (VLA) models on robots requires adapting model-specific inference pipelines to heterogeneous processors and limited onboard memory. We present vla.cpp, a unified C++ inference runtime for eleven VLA models, with no PyTorch dependency for model execution. The runtime shares model loading, tensor execution, and serving while retaining architecture-specific attention, conditioning, and action heads. Iterative policies reuse observation-dependent computation across solver steps, while regression policies predict actions directly. We evaluate task success on LIBERO-Object and profile supported configurations on NVIDIA, Apple, and Intel hardware. BitVLA completes 200/200 LIBERO-Object episodes on an 8,GB Jetson Orin Nano. A ternary tensor-core kernel accelerates its client inference by 4.0--4.6× over the CUDA-core baseline on RTX 3060 and AGX Orin. A SmolVLA case study links positional-index precision to gripper commands and task success, showing why fixed-input numerical checks should accompany rollout evaluation. Deployments on UR10e and ALOHA demonstrate physical robot integration; delay and execution-horizon studies characterize synchronous chunked control. The results demonstrate a common deployment path across VLA architectures and hardware, with numerical validation and control settings guiding deployment alongside inference efficiency.
Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen +5
VinRobotics, Vietnam · Max Planck Research School for Intelligent Systems, Germany · University of Stuttgart, Germany +3