Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration
Organizations: Shanghai Jiao Tong University
Abstract
Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself. Success rates alone cannot establish whether such gains come from better task execution or from evaluation flaws. We therefore investigate two kinds of benchmark flaws behind these gains: bugs, where the implementation does not match the intended task or evaluation protocol, and design limitations, where success criteria and simulation settings do not fully capture how acceleration affects task execution. Starting from tasks with anomalous gains, we localize root causes by plotting object trajectories against checker acceptance regions, and classify the resulting bugs into task consistency, initialization, and reproducibility. Extending this audit to seven benchmarks, including RoboTwin, LIBERO-Plus, and VLABench, we identify 22 bugs of these types and 4 design limitations. For the latter, we revise permissive success checkers, correct unrealistic object masses, and add a motion-aware score that favors smoother actions. Experiments show that bug fixes can reverse method rankings, moving the baseline from last to first on one task. Addressing design limitations can likewise remove anomalous gains: on another task, the baseline moves from 21 percentage points behind an accelerated method to 5 points ahead. Gains attributed to acceleration can therefore be artifacts of the benchmark rather than better task execution. We release our bug fixes and revised benchmark settings to support trustworthy evaluation of VLA acceleration.
Figures & tables
| Benchmark | Instruction | Demos | Evaluation |
|---|---|---|---|
| LIBERO ( Liu et al., 2023 ) | Not specified | Human teleoperation | Binary |
| LIBERO-PRO ( Zhou et al., 2025 ) | Based on LIBERO | Reused from LIBERO | Binary |
| LIBERO-Plus ( Fei et al., 2026 ) | LLM-rewritten | Reused from LIBERO | Binary |
| RoboTwin ( Chen et al., 2026b ) | LLM-generated | Scripted experts, motion planning | Binary |
| RoboCasa ( Nasiriany et al., 2024 ) | Not specified | Human teleoperation, MimicGen | Binary |
| VLABench ( Zhang et al., 2025 ) | LLM-rewritten | Scripted experts, motion planning | Binary, stage scores |
| Benchmark | Task / subset | Baseline | DP-Cache-Fast | DP-Cache-Slow | ProbeFlow |
|---|---|---|---|---|---|
| RoboTwin | place shoe | ||||
| rotate qrcode | |||||
| blocks ranking size | |||||
| handover block | |||||
| VLABench | select mahjong |
| Benchmark | Task | Baseline | DP-Cache-Fast | DP-Cache-Slow | ProbeFlow |
|---|---|---|---|---|---|
| Permissive checkers | |||||
| RoboTwin | put bottles dustbin | ||||
| VLABench | select drink | ||||
| Mass adjustment | |||||
| RoboTwin | beat block hammer | ||||
| place fan | |||||
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Repository | Commit SHA |
|---|---|
| RoboTwin | 2eeec322d95799f537cbfe5f291a8220d965ccb8 |
| LIBERO | 8f1084e3132a39270c3a13ebe37270a43ece2a01 |
| LIBERO-PRO | eafdb809426b13153aa1e4c42d6601844217dfec |
| RoboCasa365 | be22d659b02db8f6d7f3a3c3edc742934fdcbaae |
| LIBERO-Plus | 4976dc30028e805ff8094b55501d532c48fec182 |
| VLABench | cf588fe60c0c7282174fe979f5913170cfe69017 |
| Configuration | Latency (ms) | Speedup |
|---|---|---|
| Baseline | 313.73 | |
| DP-Cache-Fast | 122.61 | |
| DP-Cache-Slow | 194.12 | |
| ProbeFlow | 210.87 |
| No. | Benchmark | Bug | Repair |
|---|---|---|---|
| Task consistency — Semantic omission | |||
| 1 | RoboTwin | place shoe : the checker and demonstrations restrict toe orientation, but the instruction only specifies mat placement. | Remove the unrequested orientation constraint; retain the mat-placement check. |
| 2 | RoboTwin | rotate qrcode : the checker adds table-height and gripper-release requirements absent from the instruction. | Remove the height and release constraints; retain the orientation check. |
| 3 | RoboTwin | adjust bottle : the checker restricts lateral position by initial side, although the prompt only requests grasping and lifting. | Accept either lateral region when no arm is specified; see Figure 8 for arm-selection analysis. |
| 4 | RoboTwin | blocks ranking size : the instruction omits parts of the ordering rule used by the checker and demonstrations. | Explicitly state the required size order and placement in the instruction. |
| 5 | RoboTwin | handover block : grasp-only prompts are scored against a longer handover, placement, and release sequence. | Adapt the checker to the delivered prompt and check only its required operations. |
| Benchmark | Task | Baseline | DP-Cache-Fast | DP-Cache-Slow | ProbeFlow |
|---|---|---|---|---|---|
| Task consistency | |||||
| RoboTwin | adjust bottle | ||||
| place bread skillet | |||||
| place bread basket | |||||
| Initialization | |||||
| LIBERO-PRO | Goal-Language | ||||
| Benchmark | Task | Baseline | DP-Cache-Fast | DP-Cache-Slow | ProbeFlow |
|---|---|---|---|---|---|
| Mass adjustment | |||||
| RoboTwin | place a2b right | ||||
| stack blocks three | |||||
| LIBERO | Plate Pushing | ||||
| Motion-aware scoring | |||||
| RoboTwin | pick diverse bottles | ||||
| Task | Object | Original (g) | Adjusted (g) |
|---|---|---|---|
| RoboTwin | |||
| beat block hammer | Hammer | 1 | 600 |
| place a2b right | Bell | 50 | 150 |
| stack blocks three | Block | 10 | 50 |
| place fan | Fan | 10 | 600 |
| LIBERO | |||