LIBERO-MAX: Do Robot Policies Adapt When the World Changes?
Organizations: Tulane University · New York University · UIUC · Stanford University · CMU · University of Rochester · Nanyang Technological University · MIT CSAIL · The University of Texas at Austin
Abstract
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.
Figures & tables
| Benchmark | Static | Online | Paired | Exact | Dynamic | Scored |
|---|---|---|---|---|---|---|
| OOD | change | control | prefix | metrics | episodes | |
| LIBERO | ✗ | ✗ | ✗ | ✗ | ✗ | 2,000 |
| LIBERO-Plus | ✓ | ✗ | ✗ | ✗ | ✗ | 10,030 |
| LIBERO-PRO | ✓ | ✗ | ✗ | ✗ | ✗ | 10,000 |
| LIBERO-MAX | ✓ | ✓ | ✓ | ✓ | ✓ | 8,000 |
| Family | Model | Overall success rate (%) | Change family (paired , pp) | |||||
| Base SR | Dynamic SR | Obs. | Geom. | App. + clutter | Path | |||
| 79.7 | 65.7 | |||||||
| OpenVLA-OFT | 64.3 | 43.2 | ||||||
| X-VLA | 62.6 | 37.7 | ||||||
| Xiaomi-Robotics-0 | 70.3 | 52.0 | ||||||
| MolmoAct2 | 80.3 | 66.9 | ||||||
| Policy | Base | Reset | Mid-task |
|---|---|---|---|
| X-VLA | 68.2 | 1.3 | 12.4 |
| [65.4, 71.0] | [0.6, 2.0] | [10.4, 14.5] | |
| 79.5 | 51.3 | 55.3 | |
| [77.1, 81.9] | [48.2, 54.4] | [52.2, 58.4] | |
| HiMem-WAM | 72.5 | 68.8 | 72.6 |
| [69.7, 75.2] | [65.9, 71.6] | [69.8, 75.3] |
| Processing | Base | Dynamic | Gain |
|---|---|---|---|
| Native | 66.7 | 40.7 | — |
| Always-on | 68.3 | 47.3 | +6.7 |
| [2.3, 11.0] | |||
| Quality-gated | 67.0 | 47.7 | +7.0 |
| [3.0, 11.3] |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Categories | Events | Variants | Cases/cell | Total cases |
|---|---|---|---|---|---|
| LIBERO-Plus | 7 | 8 | 2 | 50 | 5,600 |
| LIBERO-PRO | 10 | 8 | 2 | 15 | 2,400 |
| Combined | 17 | 8 | 2 | — | 8,000 |
| Event | Family | Expected response | Frozen intervention |
|---|---|---|---|
| Target relocation | Geometry | Replan | Move the task target on its valid support after approach |
| Receptacle relocation | Geometry | Replan | Move the destination while keeping the task feasible |
| Camera shift | Observation | Continue or correct | Shift camera position, yaw, and field of view |
| Illumination switch | Appearance and clutter | Continue | Change scene lighting by a frozen scale |
| Sensor noise onset | Observation | Continue or correct | Add image corruption and a fixed occluded fraction |
| Visual theme switch | Appearance and clutter | Continue or correct | Apply one fixed color and channel transform |
| Model | Max Base | Lite Base | Max Dynamic | Lite Dynamic | Max Gap | Lite Gap |
|---|---|---|---|---|---|---|
| 79.7 | 79.3 | 65.7 | 65.1 | |||
| OpenVLA-OFT | 64.3 | 64.3 | 43.2 | 45.0 | ||
| X-VLA | 62.6 | 62.0 | 37.7 | 39.4 | ||
| Xiaomi-Robotics-0 | 70.3 | 70.6 | 52.0 | 53.5 | ||
| MolmoAct2 | 80.3 | 79.5 | 66.9 | 68.5 | ||
| SmolVLA | 26.1 | 25.6 | 15.0 | 15.3 |
| Model | Reported scale | Backbone and checkpoint training |
|---|---|---|
| Vision-language-action policies (VLA) | ||
| 2B VLM + 300M action expert | PaliGemma with a flow-matching action expert; heterogeneous robot, semantic, and web pretraining followed by LIBERO adaptation ( Black et al., 2025a ) . | |
| OpenVLA-OFT | 7B VLA + lightweight heads | Prismatic OpenVLA with continuous regression, parallel decoding, proprioception, and wrist images; combined checkpoint LoRA-optimized on the four standard LIBERO suites ( Kim et al., 2024 ; Kim et al., 2025 ) . |
| X-VLA | 0.9B | Cross-embodiment VLA with embodiment-specific soft prompts and a flow-matching decoder; released LIBERO checkpoint without MAX-specific adaptation ( Zheng et al., 2026 ) . |
| Xiaomi-Robotics-0 | 5B | Cross-embodiment and vision-language pretraining with asynchronous action-chunk deployment; released LIBERO checkpoint with its native query cadence ( Cai et al., 2026 ) . |
| MolmoAct2 | Not reported | Embodied-reasoning VLM with a flow-matching action expert conditioned through per-layer key-value caches; released LIBERO policy with continuous action decoding ( Fang et al., 2026a ) . |
| Family | Model | ||
| VLA | 50 | 5 | |
| VLA | OpenVLA-OFT | 8 | 8 |
| VLA | X-VLA | 30 | 30 |
| VLA | Xiaomi-Robotics-0 | 30 | 10 |
| VLA | MolmoAct2 | 10 | 10 |
| VLA | SmolVLA | 50 | 10 |
| Suite | Pairs | Base SR | Dynamic SR | Gap |
| LIBERO-Goal | 1,656 | 54.2 | 31.4 | |
| LIBERO-Object | 2,448 | 54.5 | 34.2 | |
| LIBERO-Spatial | 1,868 | 61.8 | 29.2 | |
| Supported total | 5,972 | 56.7 | 31.9 | |
| LIBERO-10 | 2,028 | N/A | N/A | N/A |
| Model | Preserved | Gained | Regressed | Failed | Total |
|---|---|---|---|---|---|
| Cosmos-Policy | 4,594 | 148 | 1,601 | 1,657 | 8,000 |
| 4,998 | 261 | 1,376 | 1,365 | 8,000 | |
| OpenVLA-OFT | 3,232 | 220 | 1,909 | 2,639 | 8,000 |
| X-VLA | 2,838 | 177 | 2,171 | 2,814 | 8,000 |
| Xiaomi-Robotics-0 | 3,921 | 237 | 1,703 | 2,139 | 8,000 |
| Model | Camera | Distractor | Light | Obstacle | Receptacle | Sensor | Target | Theme |
|---|---|---|---|---|---|---|---|---|
| Cosmos-Policy | ||||||||
| OpenVLA-OFT | ||||||||
| X-VLA | ||||||||
| Xiaomi-Robotics-0 | ||||||||
| MolmoAct2 |
| Model | Source | Base | Dynamic | Gap | Lowest Base | Largest loss |
|---|---|---|---|---|---|---|
| Plus | 83.0 | 63.2 | Robot initial state | Background texture | ||
| Cosmos-Policy | PRO | 64.5 | 50.2 | Position | Noise and glare | |
| Plus | 86.1 | 70.3 | Camera viewpoint | Light condition | ||
| PRO | 64.8 | 55.1 | Initial pose | Noise and glare | ||
| Plus | 68.3 | 45.2 | Robot initial state | Light condition | ||
| OpenVLA-OFT | PRO | 55.0 | 38.3 | Initial pose | Noise and glare |
| Model | Source | Cases | Base | Reset | Mid-task |
|---|---|---|---|---|---|
| X-VLA | Plus | 700 | 524 | 12 | 93 |
| PRO | 300 | 158 | 1 | 31 | |
| Plus | 700 | 606 | 382 | 411 | |
| PRO | 300 | 189 | 131 | 142 | |
| HiMem-WAM | Plus | 700 | 531 | 506 | 539 |
| PRO | 300 | 194 | 182 | 187 |
| Model | Base | Dynamic | Gap | ||
| X-VLA | 2 | 2 | 21.9 | 10.8 | |
| X-VLA | 4 | 4 | 40.0 | 19.0 | |
| X-VLA | 8 | 8 | 52.4 | 29.5 | |
| X-VLA | 12 | 12 | 57.6 | 34.6 | |
| X-VLA | 16 | 16 | 60.4 | 39.2 | |
| 2 | 10 | 74.1 | 61.1 |
| Comparison | Quantity | Estimate | 95% CI |
| Gated vs. native | Dynamic gain | [3.0, 11.3] | |
| Base gain | [ , 1.7] | ||
| Gap closure | [2.3, 11.3] | ||
| Always vs. native | Dynamic gain | [2.3, 11.0] | |
| Base gain | [ , 4.0] | ||
| Gap closure | [0.0, 10.0] |
| Group | Native | Gated | Gain | 95% CI | |
|---|---|---|---|---|---|
| Plus | 210 | 44.3 | 51.0 | [1.4, 12.4] | |
| PRO | 90 | 32.2 | 40.0 | [2.2, 13.3] | |
| Noise 24, mask 8% | 150 | 53.3 | 60.7 | [2.0, 12.7] | |
| Noise 36, mask 16% | 150 | 28.0 | 34.7 | [1.3, 12.7] |