Organizations: Tulane University · New York University · UIUC · Stanford University · CMU · University of Rochester · Nanyang Technological University · MIT CSAIL · The University of Texas at Austin
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.
Figures & tables
Figure 1 : LIBERO-MAX evaluates whether robot policies can complete tasks when the world changes during execution. Each case pairs two rollouts with identical pre-event actions: Base continues without an added event, while Dynamic receives one mid-task change. Before-and-after images illustrate four event families: geometry, observation, appearance and clutter, and path constraints. The lower panel lists all eight event types.
Figure 2 : One controlled mid-task change reduces success for every policy. We evaluate each policy on the same 8,000 pairs using its released inference settings. Arrows connect Base (open circle) to Dynamic (filled square); columns report success rates (%) and paired Dynamic minus Base changes in percentage points. Each Base/Dynamic pair shares the task, initial state, policy seed, and executed prefix, so the difference measures the effect of adding the event.
Benchmark
Static
Online
Paired
Exact
Dynamic
Scored
OOD
change
control
prefix
metrics
episodes
LIBERO
✗
✗
✗
✗
✗
2,000
LIBERO-Plus
✓
✗
✗
✗
✗
10,030
LIBERO-PRO
✓
✗
✗
✗
✗
10,000
LIBERO-MAX
✓
✓
✓
✓
✓
8,000
Table 1 : Benchmark scope and evaluation unit. ✓ marks a directly evaluated property. LIBERO scores 2,000 task episodes, LIBERO-Plus 10,030 generated variants, LIBERO-PRO 10,000 task episodes, and LIBERO-MAX 8,000 Dynamic episodes with one no-event Base control per case. Only LIBERO-MAX applies an online change after an identical executed prefix.
Figure 3 : From static source tasks to changes during execution. LIBERO-Plus and LIBERO-PRO contribute 5,600 and 2,400 source cases. Each becomes a Base/Dynamic pair with the same actions before one mid-task event. Max contains 8,000 pairs, balanced at 1,000 per event.
Family
Model
Overall success rate (%)
Change family (paired Δ , pp)
Base SR ↑
Dynamic SR ↑
Δ
Obs.
Geom.
App. + clutter
Path
π0.5
79.7
65.7
−13.9
−23.1
−24.2
−4.2
−4.2
OpenVLA-OFT
64.3
43.2
−21.1
−32.0
−33.6
−10.3
−6.8
X-VLA
62.6
37.7
−24.9
−43.0
−52.4
−1.2
−5.1
Xiaomi-Robotics-0
70.3
52.0
−18.3
−25.2
−37.7
−5.7
−3.9
MolmoAct2
80.3
66.9
−13.4
−19.3
−26.6
−4.6
−1.8
Table 2 : Every policy loses success after a controlled online change. All fourteen rows contain the same 8,000 Base/Dynamic pairs. Overall Δ and the four event-family columns report Dynamic minus Base success in percentage points. Released serving protocols differ across rows, so the table compares policies rather than isolating an architecture-family effect.
Figure 4 : Paired outcomes separate post-change regression from low Base competence. Panels partition all 8,000 pairs into the four binary outcomes. Points show episode shares; independently scaled axes retain visibility of rare change-associated successes (0,1) . The (1,0) group counts cases solved in Base but failed in Dynamic; (0,1) counts the reverse outcome.
Policy
Base
Reset
Mid-task
X-VLA
68.2
1.3
12.4
[65.4, 71.0]
[0.6, 2.0]
[10.4, 14.5]
π0.5
79.5
51.3
55.3
[77.1, 81.9]
[48.2, 54.4]
[52.2, 58.4]
HiMem-WAM
72.5
68.8
72.6
[69.7, 75.2]
[65.9, 71.6]
[69.8, 75.3]
Table 3 : Changed views can impair success from reset. SR (%) on all 1,000 cases per policy. Bold: estimates; italic below: 95% source-stratified bootstrap CIs.
Figure 5 : Geometry and observation changes produce the largest shared losses. Rows group eight events into four families. Open circles show fourteen policies per 1,000-pair event; diamonds, shading, and thin segments mark the median, interquartile range, and full range.
Figure 6 : Dynamic remains below Base at every valid query cadence. Panels use the same 800 matched pairs per policy. Open circles are Base and filled squares Dynamic; Q is the executed actions per query. Only Q≤H is valid, so π0.5 stops at Q=8 .
Processing
Base
Dynamic
Gain
Native
66.7
40.7
—
Always-on
68.3
47.3
+6.7
[2.3, 11.0]
Quality-gated
67.0
47.7
+7.0
[3.0, 11.3]
Table 4 : Restoration recovers part of the sensor-noise loss. X-VLA on 300 cases. Bold: SR (%) and Dynamic gain (pp) over Native; italic below: paired 95% bootstrap CIs for gains.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Categories
Events
Variants
Cases/cell
Total cases
LIBERO-Plus
7
8
2
50
5,600
LIBERO-PRO
10
8
2
15
2,400
Combined
17
8
2
—
8,000
Appendix
Table 5 : Outcome-independent selection balances all eight events. Cells combine source category, event, and parameter variant. Every selected case receives one Base and one Dynamic rollout, yielding 16,000 rollouts per policy.
Event
Family
Expected response
Frozen intervention
Target relocation
Geometry
Replan
Move the task target on its valid support after approach
Receptacle relocation
Geometry
Replan
Move the destination while keeping the task feasible
Camera shift
Observation
Continue or correct
Shift camera position, yaw, and field of view
Illumination switch
Appearance and clutter
Continue
Change scene lighting by a frozen scale
Sensor noise onset
Observation
Continue or correct
Add image corruption and a fixed occluded fraction
Visual theme switch
Appearance and clutter
Continue or correct
Apply one fixed color and channel transform
Appendix
Table 6 : Eight events test distinct responses to a post-commitment change. Parameters take one of two fixed variants. Continue means that task semantics remain unchanged, although the policy may still need perceptual correction; replan means that the target, destination, or path has changed.
Figure 7 : Paired observation changes. Each Base image is the exact prechange reference for the adjacent Dynamic image. The intervention changes only the named factor at the frozen trigger.
Figure 8 : Paired geometry, clutter, and path changes. Each Base image is the exact prechange reference for the adjacent Dynamic image. The task instruction remains fixed while one declared physical factor changes.
Model
Max Base
Lite Base
Max Dynamic
Lite Dynamic
Max Gap
Lite Gap
π0.5
79.7
79.3
65.7
65.1
−13.9
−14.1
OpenVLA-OFT
64.3
64.3
43.2
45.0
−21.1
−19.3
X-VLA
62.6
62.0
37.7
39.4
−24.9
−22.6
Xiaomi-Robotics-0
70.3
70.6
52.0
53.5
−18.3
−17.1
MolmoAct2
80.3
79.5
66.9
68.5
−13.4
−11.0
SmolVLA
26.1
25.6
15.0
15.3
−11.0
−10.4
Appendix
Table 7 : LIBERO-MAX Lite provides a rapid 800-pair estimate of Max on fourteen policies. Every value is success rate or Dynamic minus Base success in percentage points.
Figure 9 : The 800-pair Lite track closely reproduces Max across fourteen policies. Each policy has an upper Max track over 8,000 pairs and a lower Lite track over 800 pairs. Open circles are Base, filled squares Dynamic, and arrows connect them. The two paired changes are printed at right.
Model
Reported scale
Backbone and checkpoint training
Vision-language-action policies (VLA)
π0.5
2B VLM + 300M action expert
PaliGemma with a flow-matching action expert; heterogeneous robot, semantic, and web pretraining followed by LIBERO adaptation ( Black et al., 2025a ) .
OpenVLA-OFT
7B VLA + lightweight heads
Prismatic OpenVLA with continuous regression, parallel decoding, proprioception, and wrist images; combined checkpoint LoRA-optimized on the four standard LIBERO suites ( Kim et al., 2024 ; Kim et al., 2025 ) .
X-VLA
0.9B
Cross-embodiment VLA with embodiment-specific soft prompts and a flow-matching decoder; released LIBERO checkpoint without MAX-specific adaptation ( Zheng et al., 2026 ) .
Xiaomi-Robotics-0
5B
Cross-embodiment and vision-language pretraining with asynchronous action-chunk deployment; released LIBERO checkpoint with its native query cadence ( Cai et al., 2026 ) .
MolmoAct2
Not reported
Embodied-reasoning VLM with a flow-matching action expert conditioned through per-layer key-value caches; released LIBERO policy with continuous action decoding ( Fang et al., 2026a ) .
Appendix
Table 8 : Released checkpoints cover three policy families. Scale distinguishes reported totals from separately reported components. The final column summarizes architecture and training of the evaluated LIBERO checkpoint.
Family
Model
H
Q
VLA
π0.5
50
5
VLA
OpenVLA-OFT
8
8
VLA
X-VLA
30
30
VLA
Xiaomi-Robotics-0
30
10
VLA
MolmoAct2
10
10
VLA
SmolVLA
50
10
Appendix
Table 9 : Primary evaluations retain each policy’s serving protocol. H is the maximum number of actions returned by one query; Q is the number executed before the next observation and query.
Suite
Pairs
Base SR
Dynamic SR
Gap
LIBERO-Goal
1,656
54.2
31.4
−22.8
LIBERO-Object
2,448
54.5
34.2
−20.3
LIBERO-Spatial
1,868
61.8
29.2
−32.5
Supported total
5,972
56.7
31.9
−24.8
LIBERO-10
2,028
N/A
N/A
N/A
Appendix
Table 10 : Mimic-Video loses success across all three supported suites. Success rates are percentages, and the gap is Dynamic minus Base in percentage points. The supported total contains 5,972 pairs; the 2,028 LIBERO-10 pairs are unavailable. This diagnostic is separate from the primary 8,000-pair comparison.
Model
Preserved
Gained
Regressed
Failed
Total
(1,1)
(0,1)
(1,0)
(0,0)
Cosmos-Policy
4,594
148
1,601
1,657
8,000
π0.5
4,998
261
1,376
1,365
8,000
OpenVLA-OFT
3,232
220
1,909
2,639
8,000
X-VLA
2,838
177
2,171
2,814
8,000
Xiaomi-Robotics-0
3,921
237
1,703
2,139
8,000
Appendix
Table 11 : Paired outcomes separate post-change regression from Base competence. Counts partition the same 8,000 pairs for each of the fourteen primary policies into preserved success, change-associated success, event-associated regression, and persistent failure.
Figure 10 : Most cases reach the event and a later policy query, but many Base successes are lost. For each of fourteen policies, retained ratio is Dynamic success rate divided by Base success rate, including change-associated successes; conditional regression is the fraction of Base successes that become Dynamic failures. Trigger and response coverage each use all 8,000 evaluated cases per policy; response requires a policy query after the event. Upper filled circles show trigger coverage and lower hollow diamonds show response coverage. All values are percentages, and the three panels use separate horizontal scales.
Model
Camera
Distractor
Light
Obstacle
Receptacle
Sensor
Target
Theme
Cosmos-Policy
−23.9
−15.7
−0.3
−4.6
−24.1
−26.4
−47.5
−2.8
π0.5
−23.9
−10.9
+0.2
−4.2
−13.8
−22.3
−34.6
−2.0
OpenVLA-OFT
−35.6
−26.3
−1.2
−6.8
−25.1
−28.3
−42.1
−3.5
X-VLA
−55.9
−3.5
+0.7
−5.1
−52.4
−30.0
−52.3
−0.9
Xiaomi-Robotics-0
−30.6
−15.4
−0.1
−3.9
−31.3
−19.8
−44.0
−1.5
MolmoAct2
−17.9
−13.9
+0.6
−1.8
−21.1
−20.6
−32.0
−0.5
Appendix
Table 12 : Event sensitivity varies across policies. Entries show Dynamic minus Base success in percentage points for every policy and event. Every row has 1,000 pairs per event.
Model
Source
Base
Dynamic
Gap
Lowest Base
Largest loss
Plus
83.0
63.2
−19.8
Robot initial state
Background texture
Cosmos-Policy
PRO
64.5
50.2
−14.4
Position
Noise and glare
Plus
86.1
70.3
−15.8
Camera viewpoint
Light condition
π0.5
PRO
64.8
55.1
−9.7
Initial pose
Noise and glare
Plus
68.3
45.2
−23.0
Robot initial state
Light condition
OpenVLA-OFT
PRO
55.0
38.3
−16.7
Initial pose
Noise and glare
Appendix
Table 13 : Success falls after the change on both Plus and PRO cases. Plus contains 7 categories and 5,600 cases; PRO contains 10 categories and 2,400 cases. Base and Dynamic success rates use all assigned cases, showing where low Base success limits the size of the loss. The final columns identify the source categories with the lowest Base success and the largest loss.
Figure 11 : Paired change across all 17 source categories. (a) Category-macro Dynamic minus Base success for the seven Plus categories (circles) and ten PRO categories (squares). (b) Each model-category gap plotted against its Base success: near-zero gaps cluster where Base success is already below 20%, leaving little room for success to fall.
Figure 12 : Base and Dynamic success across the seven Plus derived source categories. Most of the 98 model-category comparisons show losses; some categories already have low Base success. Every model-category cell contains 800 pairs.
Figure 13 : Base and Dynamic success across the ten PRO derived source categories. Each category contains 240 matched pairs. Low Base success explains the small or positive differences for several pose, position, and task settings.
Figure 14 : Mid-task success exceeds Reset for all three policies. Points show paired success-rate differences, with pointwise 95% empirical bootstrap intervals, over the same 1,000 camera cases per model; resampling retains the 700 Plus / 300 PRO composition. Gray squares denote Reset minus Base, black circles Mid-task minus Base, and red diamonds Mid-task minus Reset. Gain and Loss count cases where the first condition succeeds and the second fails, or vice versa.
Model
Source
Cases
Base
Reset
Mid-task
X-VLA
Plus
700
524
12
93
PRO
300
158
1
31
π0.5
Plus
700
606
382
411
PRO
300
189
131
142
HiMem-WAM
Plus
700
531
506
539
PRO
300
194
182
187
Appendix
Table 15 : Camera-control outcomes vary across source partitions. Cells are success counts, with the source-specific denominator in the third column. The same cases appear in all three conditions.
Figure 15 : Milder camera changes improve X-VLA in both timing conditions. Reset and Mid-task use 25%, 50%, or 100% of each case’s original camera transform. The dashed Base line and pale band show the shared no-event success rate and its interval. Error bars are pointwise 95% source-stratified empirical bootstrap intervals over the same 1,000 cases. Paired Mid-task minus Reset differences (95% intervals) are +15.7 ( [12.8,18.6] ), +22.1 ( [19.3,24.9] ), and +11.1 ( [9.2,13.1] ) percentage points at 25%, 50%, and 100% strength, respectively.
Model
Q
H
Base
Dynamic
Gap
X-VLA
2
2
21.9
10.8
−11.1
X-VLA
4
4
40.0
19.0
−21.0
X-VLA
8
8
52.4
29.5
−22.9
X-VLA
12
12
57.6
34.6
−23.0
X-VLA
16
16
60.4
39.2
−21.1
π0.5
2
10
74.1
61.1
−13.0
Appendix
Table 16 : Complete results for the fixed 800-pair action cadence sweep. Base and Dynamic are success rates in percent. Gap is Dynamic minus Base in percentage points. H is the decoded action horizon and Q is the number of actions executed before the next observation. Only valid protocols with Q≤H are listed; Q∈{12,16} is therefore undefined for π0.5 at H=10 .
Comparison
Quantity
Estimate
95% CI
Gated vs. native
Dynamic gain
+7.0
[3.0, 11.3]
Base gain
+0.3
[ −1.0 , 1.7]
Gap closure
+6.7
[2.3, 11.3]
Always vs. native
Dynamic gain
+6.7
[2.3, 11.0]
Base gain
+1.7
[ −0.7 , 4.0]
Gap closure
+5.0
[0.0, 10.0]
Appendix
Table 17 : Restoration improves Dynamic success with a small aggregate Base change. Paired differences use all 300 X-VLA cases and are in percentage points. Gap closure equals Dynamic gain minus Base gain. Intervals are pointwise 95% bootstrap intervals.
Group
N
Native
Gated
Gain
95% CI
Plus
210
44.3
51.0
+6.7
[1.4, 12.4]
PRO
90
32.2
40.0
+7.8
[2.2, 13.3]
Noise 24, mask 8%
150
53.3
60.7
+7.3
[2.0, 12.7]
Noise 36, mask 16%
150
28.0
34.7
+6.7
[1.3, 12.7]
Appendix
Table 18 : Dynamic gains persist across both sources and corruption settings. Success rates are percentages; gains are percentage points for gated versus native on the same cases. Intervals are pointwise 95% bootstrap intervals. Noise denotes its standard deviation; mask denotes image occlusion.
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100% success rates, seemingly suggesting that the models are ready for deployment in real world. However, near perfect performance on existing benchmarks can be misleading: success under ideal conditions does not imply real world robustness. Existing benchmarks primarily evaluate task completion from predefined initial states, while real world interactions inevitably involve failures such as failed grasps, collisions, and unintended object movements. A robot must therefore not only execute tasks successfully, but also recognize and recover from failures to continue the task. Yet this capability remains largely unmeasured, revealing a critical gap between benchmark performance and real world reliability. To address this gap, we introduce LIBERO-Recover Benchmark, a large scale benchmark for failure recovery in robotic manipulation. Built upon LIBERO, we collect real execution failures from SOTA embodied models and construct 1,000+ scenarios across four recovery levels: (1) Action Retry, (2) Action Adaptation, (3) Object State Recovery, and (4) Environmental Recovery. We evaluate four core capabilities: spatial understanding, object structure reasoning, interaction understanding, and topological reasoning. As the first large-scale benchmark for embodied failure recovery, LIBERO-Recover shifts evaluation from \emph{Can the robot succeed?''} to \emph{Can the robot recover after failure?''}, promoting robust and generalizable embodied agents. The project will be avaible in \textcolor{blue}{https://liulin815.github.io/LIBERO-Recovery/}.
Lin Liu, Zhicheng Bao, Lu Zhang +7
Beta Infinity · Beijing Jiaotong University · School of Information and Communication Engineering Dalian University of Technology +1
Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. During closed-loop execution, actions and contacts can alter object relations and task progress, producing off-nominal intermediate states that need recovery. Recovery requires a policy to infer how task progress has changed, correct the relevant relations, and continue the original goal. We introduce RoboRecover, a benchmark for robot policy recovery under execution deviations. RoboRecover selects deviation states from trajectories, reconstructs them by replaying action prefixes, and evaluates policies on the original task. RoboRecover contains 2,000 scenarios across RoboTwin and LIBERO, with 1,000 scenarios and a fixed 800/200 train/test split on each platform. Results show that initial-state performance does not determine recovery performance and policies exhibit different recovery strengths across scenarios. Using its training split, RoboRecover further supports study on recovery interventions. RoboRecover establishes recovery from execution-induced intermediate states as a distinct dimension of robot policy evaluation.
Yang Li, Chen Zhao, Zhuoran Wang +5
School of Information, Renmin University of China, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China · University of Science and Technology of China +1
Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation task benchmarks. More recently, there has been an emphasis on evaluating the robustness of VLA models to perturbations. However, this robustness is still predominantly measured through Task Success Rate (TSR). In this work, we propose a benchmark-agnostic evaluation framework to measure the behavioural robustness of models by characterising how successful trajectories are executed under perturbation. We implement this methodology by extending the widely-used LIBERO and LIBERO-Plus benchmarks. Across three state-of-the-art VLA models, four LIBERO task suites and seven perturbation conditions, we evaluate changes in both typical successful behaviour and its variability, including metrics of motion smoothness, efficiency and gripper behaviour. We find that perturbations can alter the behaviour of successful trajectories, a phenomenon which cannot necessarily be inferred from TSR alone. Across LIBERO suites, we identify cases where state-of-the-art VLA models achieve comparable TSR under the same perturbation condition, yet behaviour on successful trajectories diverges substantially. Therefore, to have a more robust assessment of task performance, we argue that suitable measures of robustness should capture not only whether a task is completed, but also how the robot behaves while completing it. When evaluating the robustness of VLA models, TSR may be complemented by behavioural evaluation metrics that characterise the nature and variability of successful task execution by robots.
Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci +1
School of Informatics, University of Edinburgh, UK · Department of Electronics, Informatics and Bioengineering, Politecnico di Milano, Italy