cs.ROSep 29, 2026

LIBERO-MAX: Do Robot Policies Adapt When the World Changes?

Authors: Yunbei Zhang, Zijian Jin, Yuanzhe Liu, Janet Wang, Xilun Zhang, Yuyou Zhang, Zhenyu Zhang, Daoan Zhang, +9 more

Organizations: Tulane University · New York University · UIUC · Stanford University · CMU · University of Rochester · Nanyang Technological University · MIT CSAIL · The University of Texas at Austin

Abstract

Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models

    Sep 4, 2026Lin Liu, Zhicheng Bao, Lu Zhang +7Libero Manipulation BenchmarkRobotic Manipulation

  2. RoboRecover: Benchmarking Robot Policy Recovery under Execution Deviations

    Sep 24, 2026Yang Li, Chen Zhao, Zhuoran Wang +5Robot PoliciesRobot Systems

  3. Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks

    Oct 1, 2026Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci +1Libero Manipulation BenchmarkVision-Language-Action Framework