RoboRecover: Benchmarking Robot Policy Recovery under Execution Deviations
Organizations: School of Information, Renmin University of China, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China · University of Science and Technology of China · Engineering Research Center of Database and Business Intelligence, Beijing, China
Abstract
Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. During closed-loop execution, actions and contacts can alter object relations and task progress, producing off-nominal intermediate states that need recovery. Recovery requires a policy to infer how task progress has changed, correct the relevant relations, and continue the original goal. We introduce RoboRecover, a benchmark for robot policy recovery under execution deviations. RoboRecover selects deviation states from trajectories, reconstructs them by replaying action prefixes, and evaluates policies on the original task. RoboRecover contains 2,000 scenarios across RoboTwin and LIBERO, with 1,000 scenarios and a fixed 800/200 train/test split on each platform. Results show that initial-state performance does not determine recovery performance and policies exhibit different recovery strengths across scenarios. Using its training split, RoboRecover further supports study on recovery interventions. RoboRecover establishes recovery from execution-induced intermediate states as a distinct dimension of robot policy evaluation.
Figures & tables
| Setting | Recovery procedure | RSR (%) |
|---|---|---|
| Direct | Run from the recovery point | 64.4 |
| Reversion | Start 30 steps before the recovery point; run | 74.0 |
| Monitor+Reversion | Run ; at the first trigger, rollback 30 steps and resume | 77.0 |
| Correction | Run Correction for 30 steps from the recovery point; resume | 71.2 |
| Monitor+Correction | Run ; at the first trigger, apply Correction for 30 steps | 75.5 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Approach–Wrong Target Image pending Image pending RoboTwin LIBERO Redirect from an irrelevant object or region and approach the correct task target. | Approach–Wrong Pose Image pending Image pending RoboTwin LIBERO Correct the robot–object relation before attempting the intended interaction. | Approach–Stage Rollback Image pending Image pending RoboTwin LIBERO Recover lost task progress and return to a valid approach phase. |
| Approach–Scene Change Image pending Image pending RoboTwin LIBERO Reassess the changed scene and approach the currently valid object or region. | Contact–Wrong Pose Image pending Image pending RoboTwin LIBERO Restore effective contact or grasp from an unsuitable robot–object pose. | Move–Wrong Pose Image pending Image pending RoboTwin LIBERO Correct the object or end-effector pose before continuing transport. |
| Place–Wrong Target Image pending Image pending RoboTwin LIBERO Redirect the object from the wrong goal region to the instructed destination. | Place–Wrong Pose Image pending Image pending RoboTwin LIBERO Realign the object and robot before completing placement. | Place–Scene Change Image pending Image pending RoboTwin LIBERO Reassess the changed target configuration before completing placement. |
| Platform | Split | Scen. | Tasks | Natural | Pert. | Human |
|---|---|---|---|---|---|---|
| RoboTwin | Train | 800 | 42 | 800 | – | – |
| Test | 200 | 34 | 200 | – | – | |
| LIBERO | Train | 800 | 40 | 170 | 504 | 126 |
| Test | 200 | 38 | 42 | 127 | 31 |
| Policy | Initial-state [95% CI] | Recovery [95% CI] |
|---|---|---|
| RoboTwin | ||
| X-VLA | 54.33 [47.60,60.86] | 49.83 [43.25,56.35] |
| LingBot-VLA | 64.33 [57.62,70.65] | 43.20 [36.60,49.57] |
| SmolVLA | 42.67 [37.11,48.13] | 33.03 [27.66,38.65] |
| 34.50 [29.81,39.27] | 32.60 [27.48,37.83] | |
| Fast-WAM | 81.00 [75.00,86.63] | 49.33 [42.12,56.46] |
| RoboTwin policy | Macro-9 | LIBERO policy | Macro-9 |
|---|---|---|---|
| X-VLA | 40.65 | 36.79 | |
| LingBot-VLA | 41.30 | 67.71 | |
| SmolVLA | 33.52 | Being-H0.5 | 48.68 |
| 34.33 | UniFOLM | 50.26 | |
| Fast-WAM | 55.62 | Cosmos-Policy | 50.56 |
| LingBot-VA | 53.19 | Fast-WAM | 54.01 |
| Pair | Initial-state (95% CI) | Recovery (95% CI) |
|---|---|---|
| RoboTwin | ||
| -10.00 [-19.24,-0.51] | 6.63 [-3.45,16.73] | |
| 11.67 [3.85,19.50] | 16.80 [8.38,25.08] | |
| 19.83 [11.68,27.81] | 17.23 [8.16,26.21] | |
| -26.67 [-34.97,-18.26] | 0.50 [-9.02,10.23] | |
| -29.42 [-37.35,-21.49] | -9.67 [-18.77,-0.84] | |
| Model | LIBERO-Plus Total | RoboRecover Initial-state | RoboRecover RSR |
|---|---|---|---|
| 53.6 (69.4 rerun) | 85.17 | 37.80 | |
| 85.7 | 94.17 | 64.40 | |
| Being-H0.5 | 78.5 † | 91.67 | 46.80 |
| UniFOLM | – | 98.83 | 48.00 |
| Cosmos-Policy | 82.2 | 97.83 | 52.10 |
| Fast-WAM | 51.5 | 96.83 | 54.00 |
| RoboTwin Stage–Type ( ) | X-VLA | LingBot-VLA | SmolVLA |
|---|---|---|---|
| Approach–Scene Change (4) | 8.33 | 33.33 | 0.00 |
| Approach–Stage Rollback (24) | 46.67 | 26.39 | 16.11 |
| Approach–Wrong Pose (44) | 66.36 | 41.06 | 30.91 |
| Approach–Wrong Target (29) | 49.43 | 38.62 | 16.09 |
| Contact–Wrong Pose (21) | 43.81 | 65.71 | 35.87 |
| Move–Wrong Pose (35) | 53.52 | 33.33 | 46.10 |
| LIBERO Stage–Type ( ) | Being-H0.5 | ||
|---|---|---|---|
| Approach–Scene Change (28) | 12.1 | 50.0 | 24.3 |
| Approach–Stage Rollback (21) | 39.0 | 77.1 | 35.2 |
| Approach–Wrong Pose (49) | 42.4 | 59.6 | 42.4 |
| Approach–Wrong Target (25) | 40.8 | 70.4 | 58.4 |
| Contact–Wrong Pose (12) | 30.0 | 55.0 | 48.3 |
| Move–Wrong Pose (16) | 53.8 | 71.2 | 63.7 |
| Group ( ) | Being-H0.5 | ||
|---|---|---|---|
| Construction source | |||
| Human-constructed (31) | 44.5 | 70.3 | 38.7 |
| Natural (42) | 48.6 | 58.1 | 55.7 |
| Action-perturbed (127) | 32.6 | 65.0 | 45.8 |
| Task suite | |||
| LIBERO-Spatial (28) | 40.7 | 74.3 | 39.3 |
| RoboTwin source ( ) | X-VLA | LingBot-VLA | SmolVLA |
|---|---|---|---|
| X-VLA (24) | 17.78 | 42.22 | 58.06 |
| LingBot-VLA (37) | 59.10 | 19.82 | 48.83 |
| SmolVLA (83) | 43.21 | 57.67 | 17.27 |
| (56) | 67.26 | 37.62 | 35.24 |
| Split | IDs | Frames | Label 0 | Label 1 |
|---|---|---|---|---|
| Training | 720 | 320,000 | 160,000 | 160,000 |
| Balanced val. | 80 | 8,000 | 4,000 | 4,000 |
| Natural val. | 80 | 16,000 | 5,831 | 10,169 |
| Validation | Acc. | Prec. | Recall | F1 | Macro-F1 | FPR |
|---|---|---|---|---|---|---|
| Balanced | 85.49 | 86.11 | 84.63 | 85.36 | 85.49 | 13.65 |
| Natural | 86.42 | 91.99 | 86.14 | 88.97 | 85.66 | 13.09 |