Organizations: School of Information, Renmin University of China, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China · University of Science and Technology of China · Engineering Research Center of Database and Business Intelligence, Beijing, China
Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. During closed-loop execution, actions and contacts can alter object relations and task progress, producing off-nominal intermediate states that need recovery. Recovery requires a policy to infer how task progress has changed, correct the relevant relations, and continue the original goal. We introduce RoboRecover, a benchmark for robot policy recovery under execution deviations. RoboRecover selects deviation states from trajectories, reconstructs them by replaying action prefixes, and evaluates policies on the original task. RoboRecover contains 2,000 scenarios across RoboTwin and LIBERO, with 1,000 scenarios and a fixed 800/200 train/test split on each platform. Results show that initial-state performance does not determine recovery performance and policies exhibit different recovery strengths across scenarios. Using its training split, RoboRecover further supports study on recovery interventions. RoboRecover establishes recovery from execution-induced intermediate states as a distinct dimension of robot policy evaluation.
Figures & tables
Figure 1: Initial-state evaluation vs. Recovery evaluation. Recovery evaluation selects the recovery point at which an execution deviation occurs during interaction.
Figure 2: RoboRecover benchmark composition and coverage. The upper panels summarize platform, split, task, construction source, and source policies. The lower matrix shows one representative recovery point for each of the nine Stage–Type groups; cell counts report RoboTwin/LIBERO test scenarios.
Figure 3: Policy performance changes from initial-state to recovery evaluation. Lines connect each policy’s rank under matched initial-state and recovery evaluation.
Figure 4: Recovery rates across Stage–Type groups. Cells include scenario counts, and sparse groups are visually de-emphasized. Differences between groups can be hidden by Overall RSR. A·SC: Approach-Scene Change.
Figure 5: Cross-policy recovery on Natural deviations. Rows identify source policies, and columns show the evaluated recovery policies. Solid outlines mark source-policy re-evaluation; dashed outlines mark the observed-best alternative.
Setting
Recovery procedure
RSR (%)
Direct π0.5
Run π0.5 from the recovery point
64.4
Reversion
Start 30 steps before the recovery point; run π0.5
74.0
Monitor+Reversion
Run π0.5 ; at the first trigger, rollback 30 steps and resume
77.0
Correction
Run Correction for 30 steps from the recovery point; resume π0.5
71.2
Monitor+Correction
Run π0.5 ; at the first trigger, apply Correction for 30 steps
75.5
Table 1: LIBERO recovery-intervention results. Reversion changes the startingstate, whereas Correction adds a 30-step learned controller. Monitor-basedsettings activate the corresponding intervention selectively.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Approach–Wrong Target Image pending Image pending RoboTwin LIBERO Redirect from an irrelevant object or region and approach the correct task target.
Approach–Wrong Pose Image pending Image pending RoboTwin LIBERO Correct the robot–object relation before attempting the intended interaction.
Approach–Stage Rollback Image pending Image pending RoboTwin LIBERO Recover lost task progress and return to a valid approach phase.
Approach–Scene Change Image pending Image pending RoboTwin LIBERO Reassess the changed scene and approach the currently valid object or region.
Contact–Wrong Pose Image pending Image pending RoboTwin LIBERO Restore effective contact or grasp from an unsuitable robot–object pose.
Move–Wrong Pose Image pending Image pending RoboTwin LIBERO Correct the object or end-effector pose before continuing transport.
Place–Wrong Target Image pending Image pending RoboTwin LIBERO Redirect the object from the wrong goal region to the instructed destination.
Place–Wrong Pose Image pending Image pending RoboTwin LIBERO Realign the object and robot before completing placement.
Place–Scene Change Image pending Image pending RoboTwin LIBERO Reassess the changed target configuration before completing placement.
Appendix
Figure 6: Representative recovery observations for the nine occupied Stage–Type groups. Each panel pairs RoboTwin (left) and LIBERO (right) examples and summarizes the recovery required to continue the original task.
Platform
Split
Scen.
Tasks
Natural
Pert.
Human
RoboTwin
Train
800
42
800
–
–
Test
200
34
200
–
–
LIBERO
Train
800
40
170
504
126
Test
200
38
42
127
31
Appendix
Table 3: Corpus composition. Pert. denotes Action-Perturbed scenarios; dashes indicate construction sources not used on RoboTwin.
Policy
Initial-state [95% CI]
Recovery [95% CI]
RoboTwin
X-VLA
54.33 [47.60,60.86]
49.83 [43.25,56.35]
LingBot-VLA
64.33 [57.62,70.65]
43.20 [36.60,49.57]
SmolVLA
42.67 [37.11,48.13]
33.03 [27.66,38.65]
π0.5
34.50 [29.81,39.27]
32.60 [27.48,37.83]
Fast-WAM
81.00 [75.00,86.63]
49.33 [42.12,56.46]
Appendix
Table 4: Complete initial-state and recovery results on the fixed test splits (%). Bold marks the highest point estimate in each platform and evaluation condition.
RoboTwin policy
Macro-9
LIBERO policy
Macro-9
X-VLA
40.65
π0
36.79
LingBot-VLA
41.30
π0.5
67.71
SmolVLA
33.52
Being-H0.5
48.68
π0.5
34.33
UniFOLM
50.26
Fast-WAM
55.62
Cosmos-Policy
50.56
LingBot-VA
53.19
Fast-WAM
54.01
Appendix
Table 5: Macro-9 RSR on the fixed test splits (%), computed from unrounded scenario-level rates. Best values are bold.
Pair A−B
Initial-state Δ (95% CI)
Recovery Δ (95% CI)
RoboTwin
X−L
-10.00 [-19.24,-0.51]
6.63 [-3.45,16.73]
X−S
11.67 [3.85,19.50]
16.80 [8.38,25.08]
X−P
19.83 [11.68,27.81]
17.23 [8.16,26.21]
X−F
-26.67 [-34.97,-18.26]
0.50 [-9.02,10.23]
X−V
-29.42 [-37.35,-21.49]
-9.67 [-18.77,-0.84]
Appendix
Table 6: All paired policy contrasts in percentage points. RoboTwin intervals use trajectory-cluster bootstrap; LIBERO intervals use paired scenario bootstrap.
Model
LIBERO-Plus Total
RoboRecover Initial-state
RoboRecover RSR
π0
53.6 (69.4 rerun)
85.17
37.80
π0.5
85.7
94.17
64.40
Being-H0.5
78.5 †
91.67
46.80
UniFOLM
–
98.83
48.00
Cosmos-Policy
82.2
97.83
52.10
Fast-WAM
51.5
96.83
54.00
Appendix
Table 7: Published LIBERO-Plus results and RoboRecover results for overlapping model families (%). The π0 value in parentheses is the JAX rerun reported by Zhang et al. (2026) . † Being-H0.5 is reported by Luo et al. (2026b) . ‡ SmolVLA is reported under a separate LIBERO training and evaluation setup by Li et al. (2026a) . Dashes indicate that no directly reported aggregate LIBERO-Plus result was identified.
RoboTwin Stage–Type ( N )
X-VLA
LingBot-VLA
SmolVLA
Approach–Scene Change (4)
8.33
33.33
0.00
Approach–Stage Rollback (24)
46.67
26.39
16.11
Approach–Wrong Pose (44)
66.36
41.06
30.91
Approach–Wrong Target (29)
49.43
38.62
16.09
Contact–Wrong Pose (21)
43.81
65.71
35.87
Move–Wrong Pose (35)
53.52
33.33
46.10
Appendix
Table 8: Complete RoboTwin recovery rates across Stage–Type groups (%). The two Scene Change groups contain only four and two scenarios.
LIBERO Stage–Type ( N )
π0
π0.5
Being-H0.5
Approach–Scene Change (28)
12.1
50.0
24.3
Approach–Stage Rollback (21)
39.0
77.1
35.2
Approach–Wrong Pose (49)
42.4
59.6
42.4
Approach–Wrong Target (25)
40.8
70.4
58.4
Contact–Wrong Pose (12)
30.0
55.0
48.3
Move–Wrong Pose (16)
53.8
71.2
63.7
Appendix
Table 9: Complete LIBERO recovery rates across Stage–Type groups (%). Cosmos denotes Cosmos-Policy. The Place–Wrong Target group contains seven scenarios and is reported descriptively.
Group ( N )
π0
π0.5
Being-H0.5
Construction source
Human-constructed (31)
44.5
70.3
38.7
Natural (42)
48.6
58.1
55.7
Action-perturbed (127)
32.6
65.0
45.8
Task suite
LIBERO-Spatial (28)
40.7
74.3
39.3
Appendix
Table 10: LIBERO recovery rates by construction source and task suite (%).
RoboTwin source ( N )
X-VLA
LingBot-VLA
SmolVLA
X-VLA (24)
17.78
42.22
58.06
LingBot-VLA (37)
59.10
19.82
48.83
SmolVLA (83)
43.21
57.67
17.27
π0.5 (56)
67.26
37.62
35.24
Appendix
Table 11: Complete Natural source-policy by recovery-policy matrices (%). Cosmos denotes Cosmos-Policy. Rows with small N are descriptive.
Split
IDs
Frames
Label 0
Label 1
Training
720
320,000
160,000
160,000
Balanced val.
80
8,000
4,000
4,000
Natural val.
80
16,000
5,831
10,169
Appendix
Table 12: RoboTwin monitor training and validation data. The two validation sets use different frame-sampling distributions from the same 80 held-out scenario IDs.
Validation
Acc.
Prec.
Recall
F1
Macro-F1
FPR
Balanced
85.49
86.11
84.63
85.36
85.49
13.65
Natural
86.42
91.99
86.14
88.97
85.66
13.09
Appendix
Table 13: RoboTwin monitor validation results (%). The lower panel gives the corresponding confusion matrices. FPR is the false-positive rate.
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100% success rates, seemingly suggesting that the models are ready for deployment in real world. However, near perfect performance on existing benchmarks can be misleading: success under ideal conditions does not imply real world robustness. Existing benchmarks primarily evaluate task completion from predefined initial states, while real world interactions inevitably involve failures such as failed grasps, collisions, and unintended object movements. A robot must therefore not only execute tasks successfully, but also recognize and recover from failures to continue the task. Yet this capability remains largely unmeasured, revealing a critical gap between benchmark performance and real world reliability. To address this gap, we introduce LIBERO-Recover Benchmark, a large scale benchmark for failure recovery in robotic manipulation. Built upon LIBERO, we collect real execution failures from SOTA embodied models and construct 1,000+ scenarios across four recovery levels: (1) Action Retry, (2) Action Adaptation, (3) Object State Recovery, and (4) Environmental Recovery. We evaluate four core capabilities: spatial understanding, object structure reasoning, interaction understanding, and topological reasoning. As the first large-scale benchmark for embodied failure recovery, LIBERO-Recover shifts evaluation from \emph{Can the robot succeed?''} to \emph{Can the robot recover after failure?''}, promoting robust and generalizable embodied agents. The project will be avaible in \textcolor{blue}{https://liulin815.github.io/LIBERO-Recovery/}.
Lin Liu, Zhicheng Bao, Lu Zhang +7
Beta Infinity · Beijing Jiaotong University · School of Information and Communication Engineering Dalian University of Technology +1
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.
Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many existing recovery methods rely on human intervention. In this paper, we propose Failure-Aware Retry (FAR), a framework that enables robots to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete the task autonomously. FAR combines Failure-Contrastive Preference Adaptation, which constructs preference learning data from failures to steer the policy away from previously unsuccessful behaviors, with lightweight action perturbations during retries to encourage local exploration. We further incorporate successful recovery trajectories into a training loop for continual policy improvement. Experiments in both simulation and real-world manipulation tasks show that FAR substantially improves success rates and robustness, with average gains of 17.6% over the standard diffusion policy in simulation and 11.7% in the real world. In addition, FAR significantly improves data efficiency under both reset and timestep budgets during continual policy improvement by exploiting informative failure cases.