Manipulation failures can leave scenes in states from which a task policy cannot recover. Learning corrective behaviors requires scalable failure exploration and physical grounding. We present Recova, an agent-guided framework that jointly develops task execution and recovery in a reconstructed digital twin, then verifies and refines both through real-world experience. In the twin, the agent diagnoses failures, tests corrective programs, and collects successful task and recovery rollouts for separate policies. During deployment, it monitors progress, invokes a learned or programmatic recovery, verifies scene restoration, and resumes execution. When no suitable recovery is available, a human demonstration resolves the failure and enters the learning loop, allowing the system to expand its recovery capabilities. Physical rollouts and human demonstrations are routed to the corresponding policy for DAgger training. Across six LIBERO-Pro settings and four MolmoSpaces categories, Recova achieves 78.8% and 64.9% mean success, compared with 71.7% and 38.0% for the strongest baselines. With parallel collection across four real-robot workstations, DAgger fine-tuning raises mean success from 23.8% to 77.5%, and recovery skills further raise it to 87.5%. Over four collection rounds on one task, observed human intervention falls from 87.5% to 0%. Together, these results show how agent-guided recovery turns failures into reusable capabilities, improving robustness while progressively reducing human intervention. Project page: https://www.liuisabella.com/Recova
Figures & tables
Figure 2: (A) Simulation: Recova builds a digital twin from camera observations, calibration, and robot trajectories. Successful task and recovery rollouts train separate policies, while recovery programs populate a code-as-policy skill library. (B) Real-robot rollout: Recova monitors execution, invokes recovery policies or programs, verifies restoration, and requests human demonstrations when needed. Task successes, verified recoveries, and human demonstrations update the corresponding policies, while new failures guide skill-library expansion. Paired frames show rollout endpoints; the inset tracks human intervention across DAgger rounds.
Figure 3: Recovery skills in simulation. The failure (left) and recovery (right) moments of one LIBERO-Pro (top) or MolmoSpaces (bottom) episode: the main camera view, with the wrist camera inset at the top right, under the recovery instruction.
MolmoSpaces ↑
Methods
Pick
P&P
Open
Close
Avg.
Agentic code-as-policy methods
CaP-Agent0 ( 16 )
23.0
11.0
14.0
36.0
21.0
RATs ( 67 )
37.0
22.0
20.0
73.0
38.0
Vision-language-action policies
π0 ( 4 )
14.0
7.0
10.0
47.0
19.5
Table 1: MolmoSpaces task success (%). P&P denotes pick-and-place. Avg. is the unweighted mean across the four simulation task categories. Bold marks the best reported value in each column. Baseline rows reproduce published results.
Object ↑
Goal ↑
Spatial ↑
Methods
Pos.
Task
Pos.
Task
Pos.
Task
Avg. ↑
Agentic code-as-policy methods
CaP-Agent0 ( 16 )
27.0
31.0
29.0
16.0
13.0
23.0
23.2
RATs ( 67 )
61.0
63.0
43.0
36.0
29.0
31.0
43.8
ASPIRE ( 41 )
98.0
95.0
81.0
45.0
51.0
60.0
71.7
Vision-language-action policies
Table 2: LIBERO-Pro task success (%). Object, Goal, and Spatial each include initial-position swaps (Pos.) and task perturbations (Task). Avg. is the unweighted mean of all six columns. Bold marks the best reported value in each column. Baseline rows reproduce published results.
Figure 4: Parallel automated DAgger data collection. Left: Four robot stations operate in parallel with varied scene configurations. Middle: Station activity during one collection session, distinguishing task-policy execution, recovery-policy execution, human control, and waiting or reset periods. Right: Observed task success and human intervention rates across four collection rounds. As policies are fine-tuned on accumulated data, task success increases while human intervention decreases.
Policy
Pencil box
Stack rings
Draw tile
Discard tile
Avg. ↑
Base task policy
20.0
35.0
25.0
15.0
23.8
+ DAgger
70.0
75.0
85.0
80.0
77.5
+ Recovery skills
85.0
90.0
90.0
85.0
87.5
Table 3: Real-robot task success (%) over 20 trials per task and configuration. Rows add DAgger fine-tuning of the task policy, then recovery skills at deployment. Bold marks the best result for each task.
Figure 5: Real-robot tasks and their recovery skills. For each task, the gray panel shows the start and result of a task-policy rollout. Each recovery skill appears under its name as two frames of one rollout, the action and then the result, developed in the digital twin (blue) or executed on the robot by a learned recovery policy (green). The list on the right names further recovery skills.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Monitor prompts , abridged: each query’s attached text and images, prompt, and required JSON reply. Symbols follow Alg. 1 .
Figure 7: Recorded monitor verdicts while drawing a mahjong tile. Each image is the four-camera composite sent to the monitor. (a) The monitor reuses a known recovery skill verbatim. (b) A stalled task in an intact scene leads to a task demonstration. (c) The recovery policy executes the proposed instruction, and the restoration query confirms the scene. (d) The completion query accepts a finished task.
Task
Success criterion
Pencil box
The pencil lies in its tray, and the tray is slid back into the box sleeve.
Stack rings
The four rings sit on the peg in size order.
Draw tile
Exactly one more tile stands upright at the end of the hand row, facing the front camera; no tile has fallen, and the tile wall is otherwise unchanged.
Discard tile
The specified hand tile lies face-up in the central discard area, the other hand tiles are undisturbed, and both grippers are empty.
Table 5: Policy fine-tuning hyperparameters , shared by all task and recovery policies.
Figure 8: Simulated recovery skills for pencil-box packing and ring stacking. Each blue panel shows one recovery skill developed in the task’s digital twin, as four frames of one rollout, from the failure to the result. Frames come from the twin’s inspection camera, except Move the idle arm clear (top camera).
Figure 9: Simulated recovery skills for drawing and discarding a mahjong tile. Panels follow the conventions of Fig. 8 .
Figure 10: Recovery examples on LIBERO-Pro (top) and MolmoSpaces (bottom). Each panel shows one episode’s failure and recovery moments under its recovery instruction and task, as in Fig. 3 . Red rings in some LIBERO-Pro frames are markers from the source recordings.
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100% success rates, seemingly suggesting that the models are ready for deployment in real world. However, near perfect performance on existing benchmarks can be misleading: success under ideal conditions does not imply real world robustness. Existing benchmarks primarily evaluate task completion from predefined initial states, while real world interactions inevitably involve failures such as failed grasps, collisions, and unintended object movements. A robot must therefore not only execute tasks successfully, but also recognize and recover from failures to continue the task. Yet this capability remains largely unmeasured, revealing a critical gap between benchmark performance and real world reliability. To address this gap, we introduce LIBERO-Recover Benchmark, a large scale benchmark for failure recovery in robotic manipulation. Built upon LIBERO, we collect real execution failures from SOTA embodied models and construct 1,000+ scenarios across four recovery levels: (1) Action Retry, (2) Action Adaptation, (3) Object State Recovery, and (4) Environmental Recovery. We evaluate four core capabilities: spatial understanding, object structure reasoning, interaction understanding, and topological reasoning. As the first large-scale benchmark for embodied failure recovery, LIBERO-Recover shifts evaluation from \emph{Can the robot succeed?''} to \emph{Can the robot recover after failure?''}, promoting robust and generalizable embodied agents. The project will be avaible in \textcolor{blue}{https://liulin815.github.io/LIBERO-Recovery/}.
Lin Liu, Zhicheng Bao, Lu Zhang +7
Beta Infinity · Beijing Jiaotong University · School of Information and Communication Engineering Dalian University of Technology +1
Robust embodied robots should be able to recover from failures and retry tasks in order to operate reliably in unstructured and noisy real-world environments. Achieving this capability requires training policies on data that captures recovery behaviors. However, collecting such data through robot teleoperation is difficult to scale, as it is time-consuming to induce diverse failure states, perform corrective actions, and reset the environment. This challenge is further exacerbated by the high diversity of failure modes, which demands substantially more recovery data than success demonstrations. In this work, we show that egocentric human data capturing failure recovery processes provides a scalable alternative. By efficiently arranging task-level failure configurations and recording short recovery segments, human operators can generate more than 10x as much valid recovery data per hour compared to robot teleoperation under our protocol. To address the embodiment gap between human and robot, we propose EgoRecovery, a co-training framework for learning recovery behavior, where human recovery demonstrations are aligned to a compact corrective-intent space shared with robot data, which captures the timing and magnitude of correction. Only a small number of robot recovery demonstrations are required to connect this intent to executable robot actions. At deployment, a learned recovery gate predicts when correction is needed from robot observations and activates the corrective intent only in recovery states. Experiments on real-world recovery tasks show that EgoRecovery improves success from failure starts over robot-only recovery, direct co-training with human recovery data, and direct intent-transfer baselines.
Zuhao Ge, Yuchen Zhou, Weitao Zhou +8
Institute of Trustworthy Embodied AI (TEAI), Fudan University · 2Shanghai Key Laboratory of Multimodal Embodied AI · 3Simple AI
Hierarchical robotic systems executing long-horizon manipulation tasks must make high-level semantic decisions that orchestrate stochastic low-level skills. In this setting, failed rollouts are ambiguous: a poor downstream state may reflect an invalid high-level decision, partial observation, or a valid decision whose physical execution failed. Traditional supervised learning lacks data for such recovery states, while reinforcement learning struggles with sparse rewards and non-local credit assignment. We propose MAGMA-GEN, an on-policy data-generation pipeline that converts ambiguous failed rollouts into validated recovery supervision. MAGMA-GEN first uses a privileged coach to hypothesize an early decision-level error and propose localized correction or recovery actions. Because this diagnosis is fallible, candidates are retained only if re-execution from the same state under matched conditions improves downstream progress. This produces supervised examples from the agent's own failure distribution without per-step human demonstrations. Evaluated on interactive long-horizon manipulation tasks, MAGMA-GEN improves task success and recovery capabilities, against distillation and trajectory-repair baselines under evolving task constraints in both simulation and real-robot execution.