RESETTLE: Robotic Recovery through Disagreement-Triggered Retrieval and Efficient Corrective Control
Authors: Yuxin Chen, Senqiao Yang, Zixuan Wang, Jinhui Ye, Changsheng Lu, Pengguang Chen, Shu Liu, Zhuotao Tian, +1 more
Organizations: The Hong Kong University of Science and Technology · The Chinese University of Hong Kong · SmartMore · Harbin Institute of Technology, Shenzhen · Shenzhen Loop Area Institute
Reliable robotic manipulation requires timely intervention to correct emerging deviations and restore progress after execution errors. However, recovery methods based on repeated vision-language reasoning or iterative online optimization can incur substantial latency, delaying intervention. To address these challenges, we introduce RESETTLE(Robotic rEcovery through diSagrEement-Triggered reTrievaL and Efficient Corrective Control), a model-agnostic framework that provides computationally efficient recovery at the action-execution interface of frozen robot policies. RESETTLE triggers recovery when two action proposals independently sampled under identical conditioning persistently disagree. It retrieves a same-task demonstration reference using an adapted V-JEPA encoder and combines a state-servo prior with a guarded visual residual to execute one corrective action without online trajectory optimization or additional vision-language reasoning, then returns control to the base policy. Across six base policies in simulation, RESETTLE achieves up to 8.70%, 6.28%, and 6.83% absolute success-rate gains on LIBERO-Plus, Meta-World, and RoboCasa Tabletop, respectively, with further improvements on four real-world tasks using two policies. In QwenPI-based comparisons, its monitoring-and-recovery computation latency is 74.04%--93.57% lower than VoLoAgent's monitoring-and-planning latency for grasp and place tool calls. It also raises Harness VLA's LIBERO-Pro Swap success from 42% to 50%, demonstrating compatibility with high-level agentic planning. Code available at: https://github.com/JIA-Lab-research/RESETTLE
Figures & tables
Figure 1: QwenPI action disagreement on LIBERO-Plus for perturbations in layout (left) and robot-initialization (right) . Histograms show within-group decision percentages. “Training Data” denotes training-data inference; “Training P95” marks its 95th percentile. Failures exhibit heavier tails.
Figure 2: Overview of RESETTLE. Disagreement-Based Intervention module detects the execution error and initiates V-JEPA-Based Retrieval, followed by Reference-Guided Recovery using a state-servo prior and a guarded visual residual. The a1,a2 are action chunks sampled with independent seeds; Ocurr/ref and Scurr/ref denote current/reference observations and measured robot states. Each intervention executes one native control tick before re-observation and return to the base policy. ‘Snowflake’ signifies frozen policy while ‘flame’ refers to offline encoder adaptation.
Figure 3: Experimental settings across four simulation benchmarks and four real-world manipulation tasks using a Franka Research 3 robot with external and wrist-mounted cameras.
Model
Cam.
Robot
Lang.
Light
BG
Noise
Layout
Average
OpenVLA
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
OpenVLA-OFT
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
NORA
2.2
37.0
65.1
45.7
58.6
12.8
62.1
39.0
WorldVLA
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
UniVLA
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
π0
13.8
6.0
58.8
85.0
81.4
79.0
68.8
53.6
Table 1: Results on LIBERO-Plus. Per-dimension success rates (%) on the LIBERO-Plus robustness benchmark. Cam.: camera; Lang.: language; BG: background; Robot: robot initialization. All models are trained on standard LIBERO and evaluated on the perturbed test set. Average aggregates all evaluation cases, weighting each dimension by its number of cases. “Ours” denotes RESETTLE, while Ours † denotes RESETTLE without the residual module.
Model
Easy (28)
Medium (11)
Hard (6)
Very Hard (5)
Average
QwenPI / Base.
80.36
57.27
45.00
60.00
60.66
QwenPI / Ours
83.57 3.21↑
62.73 5.46↑
51.67 6.67↑
66.00 6.00↑
65.99 5.33↑
SmolVLA / Base.
79.64
41.82
83.33
32.00
59.20
SmolVLA / Ours
83.57 3.93↑
46.36 4.54↑
90.00 6.67↑
42.00 10.00↑
65.48 6.28↑
Table 2: Success rates (%) across difficulty levels on Meta-World. Parentheses indicate the number of tasks in each difficulty group. “Ours” denotes RESETTLE. Arrows indicate absolute improvements (%) over the corresponding baseline.
Table 6
Model
Press-button
Cube-up
Carrot-in-pot
Use Spoon
Avg.
SR
SR
SR
SC
QwenPI / Base.
90
55
85
36.7
66.67
QwenPI / Ours
100
70
95
50
78.75
VLAct / Base.
90
85
90
80
86.25
VLAct / Ours
100
90
95
86.7
92.92
Table 5: Real-world performance. Performance (%) over 20 trials per task and method. Press-button, Cube-up, and Carrot-in-pot report success rate (SR), while Use Spoon reports normalized stage completion (SC). “Ours” denotes RESETTLE. Avg. is the mean across four tasks.
Method
Monitor
Recovery
Total ↓
Relative latency ↓
RESETTLE (Ours)
≈ 0.0
83.6
≈ 83.6
1.0×
Agentic Robot
107.0
109.0
216.0
2.6×
VoLoAgent
107.0
215.0–1193.0
322.0–1300.0
3.9 – 15.6×
LPB †
56.2
423.5
479.7
5.7×
Table 6: Monitoring and recovery computation latency. Monitor, Recovery, and Total are in milliseconds; lower is better. Total is Monitor + Recovery under the measured QwenPI-based configurations. Monitor denotes additional latency beyond nominal policy inference. Relative latency is Total divided by RESETTLE’s 83.6 ms. VoLoAgent ranges give the place and grasp branches as lower and upper endpoints, respectively. LPB † is implemented in our QwenPI+V-JEPA framework.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Paired grasp and contact recovery cases. QwenPI baseline and QwenPI+RESETTLE use the same task, layout, and seed in each pair. Panels (a)–(d) show grasp-pose correction, regrasping after a drop, recovery from an obstructed grasp, and correction during stove-knob operation. The upper yellow sequence is the baseline; the lower sequence alternates green base-policy execution with red RESETTLE interventions. Red frames labeled “VJEPA” denote the recovery controller using V-JEPA representations, reference retrieval, and guarded residual control. Frames progress from left to right.
Figure 5: Paired placement and grasp correction cases. Panels (a)–(c) show recovery during obstructed cup placement, correction after unintended object grasping, and grasp-pose adjustment that avoids an unfavorable dropped-object state. Each pair matches the task, layout, and seed, with yellow baseline frames above and green base-policy/red recovery frames below, following Figure 4 . Panel (c) includes recovery at two stages, interleaved with base-policy execution. Selected frames show local corrections and subsequent progress rather than every control step.
Figure 6: Effect of calibration set size on action disagreement. QwenPI distributions from 5, 100, and all training trajectories across seven LIBERO-Plus perturbation categories. The similar distribution shapes support lightweight threshold calibration using only five training trajectories.
Module
Total parameters
Trainable parameters
Context encoder
1,012,173,952
1,012,173,952
EMA target encoder
1,012,173,952
0
Action-conditioned predictor
305,220,992
305,213,824
Guarded residual-action head
6,839,226
6,839,226
All instantiated modules
2,336,408,122
1,324,227,002
Deployed subset
1,019,013,178
–
Appendix
Table 7: Parameter counts for the LIBERO recovery module. Trainable counts include only parameters optimized by backpropagation. The deployed subset consists of the context encoder and residual head; its trainable count is omitted because no optimization occurs during deployment. Counts exclude the base policy, reference-bank features, optimizer states, and non-parameter buffers.
Setting
LIBERO
RoboCasa
Meta-World
State/action dim.
7
29
4
Frames per clip
8
8
8
Input resolution
2562
2562
2562
Per-GPU batch
16
16
16
Global batch
128
128
128
Training budget
20k updates
150k updates
20 epochs
Appendix
Table 8: Recovery-model training budgets. Training budgets for the benchmark-specific recovery models. Batch sizes count clips. A virtual epoch denotes a fixed number of optimizer updates rather than a complete dataset traversal.
Figure 7: Feature-distance and action-disagreement distributions across six LIBERO-Plus perturbation categories: language instructions, background textures, camera viewpoints, robot initial states, layout perturbations, and light conditions. Each pair shows V-JEPA distance to same-task training references (left, frame-level within-outcome percentages) and QwenPI action disagreement (right, within-group policy-decision percentages). Red and blue denote failed and successful executions; green in disagreement panels denotes offline training-data inference. n counts frames or policy decisions in the respective panels, not rollouts. Vertical lines mark “Success P95” for feature distance and “Training P95” for action disagreement.
Perturbation
V-JEPA Feature Distance
Action Disagreement (Ours)
Failed
Successful
Failed
Successful
Rate (%)
Count
Rate (%)
Count
Rate (%)
Count
Rate (%)
Count
Camera
26.52
179.3
10.00
66.9
93.10
16.6
29.47
2.4
Robot
27.98
104.6
17.44
28.2
95.40
16.9
25.40
2.6
Language
15.48
70.5
17.89
24.8
77.89
9.6
9.22
1.9
Light
21.74
10.2
9.74
59.4
91.67
12.6
12.34
1.8
Appendix
Table 9: Failure coverage under two consecutive P95 exceedances. Rate: trajectories triggered (%). Count: mean triggers per triggered trajectory.
Setting
Failed
Successful
Rate (%)
Count
Rate (%)
Count
FastWAM / LIBERO-Plus
Camera
42.01
2.0
12.99
1.5
Robot
53.36
2.4
11.19
1.4
Language
68.99
3.5
8.91
1.4
Light
64.83
3.0
10.71
1.2
Appendix
Table 10: Action-disagreement triggers for additional policies and benchmarks, using training-data inference P95 thresholds and two consecutive exceedances with reset after each trigger. Rate: trajectories triggered (%). Count: mean triggers per triggered trajectory.
Perturbation
Mean takeovers per trajectory
Intervened trajectories (%)
Recovery actions (%)
All
Intervened
Camera
3.90
8.80
44.28
1.62
Robot
25.32
29.73
85.16
12.59
Language
0.07
2.78
2.34
0.04
Light
0.07
2.58
2.71
0.05
Background
0.29
2.70
10.78
0.19
Appendix
Table 11: Actual recovery usage for QwenPI+RESETTLE on LIBERO-Plus using a shared training-data inference P95 threshold and two consecutive exceedances. Each takeover executes one native recovery action. Mean counts include either all trajectories or only those with at least one takeover (Intervened). Recovery actions (%) denotes the share of all executed actions, excluding initialization waiting actions. Overall statistics are computed from pooled counts.
Setting
Cam.
Robot
Lang.
Light
BG
Noise
Layout
Total
Baseline
55.7
61.0
94.5
98.0
97.6
86.2
79.1
80.2
P85
58.7
84.2
94.7
97.5
97.7
79.5
78.8
83.1
P90
59.2
83.8
94.3
98.0
97.9
83.3
79.3
83.9
P95
59.9
82.4
94.7
97.8
98.6
86.9
80.1
84.6
P99
57.3
76.2
94.9
97.5
98.5
89.6
80.1
83.6
Appendix
Table 12: Threshold percentile ablation on LIBERO-Plus. Success rates (%) for QwenPI+RESETTLE using training-data inference thresholds and two consecutive exceedances. Baseline denotes QwenPI without RESETTLE, as reported in Table 1 . Total aggregates all evaluation cases, weighting each perturbation category by its number of cases. Cam.: camera; Lang.: language; BG: background; Robot: robot initialization. Values are rounded to one decimal place; bold marks the best result in each column, including ties.
Figure 8: Failures without a recovery trigger. Seven unsuccessful rollouts with no monitoring trigger, ordered from top to bottom. Frames progress left to right. The cases include contact-induced scene changes, misplaced objects, obstructed motions, and incorrect grasp or task decisions. Consistent policy proposals can therefore accompany unsuccessful execution.
Figure 9: Failures despite triggered recovery. Seven unsuccessful rollouts in which recovery is triggered, ordered from top to bottom, with frames progressing left to right. Cases illustrate contact-induced deterioration, unintended grasp configurations, execution-budget exhaustion, and task-goal mismatch. Triggered intervention does not guarantee completion within the evaluation budget.
Figure 10: Real-world training demonstrations. Rows show Carrot-in-pot, Press-button, Cube-up, and Use Spoon, from top to bottom, with time progressing from left to right. Selected frames illustrate object acquisition and placement, button contact, stacking, and tool use. Successful training demonstrations support recovery-module learning and provide same-task references for retrieval.
Figure 11: Real-world execution with QwenPI+RESETTLE. Task groups show Carrot-in-pot, Press-button, Cube-up, and Use Spoon from top to bottom. Frames progress left to right; Cube-up continues from its upper row to its lower row within one rollout. All sequences use QwenPI+RESETTLE, and individual recovery actions are not annotated. The selected frames show the combined system’s execution, including repeated approaches before cube acquisition and subsequent stacking.
ID
Task
QwenPI
SmolVLA
Base.
Ours
Base.
Ours
00
nut-assembly
40
40
90
90
01
basketball
30
30
10
10
02
bin-picking
40
40
10
20
03
box-close
30
70
40
50
04
button-press-topdown
100
100
100
100
Appendix
Table 13: Per-task ASR on Meta-World (tasks 00–24). Success rates (%) over 10 episodes per task. Ours denotes RESETTLE; bold marks the better result within each model pair, including ties.
ID
Task
QwenPI
SmolVLA
Base.
Ours
Base.
Ours
25
handle-pull-side
10
20
20
30
26
handle-pull
10
30
100
100
27
lever-pull
40
40
0
10
28
pick-place-wall
70
70
40
20
29
pick-out-of-hole
0
30
40
30
Appendix
Table 14: Per-task ASR on Meta-World (tasks 25–49). Success rates (%) over 10 episodes per task. Ours denotes RESETTLE; bold marks the better result within each model pair, including ties.
ID
Task
QwenGR00T
LDA
Base.
Ours
Base.
Ours
01
Cup → Drawer (close)
36
34
50
42
02
Potato → Microwave (close)
30
42
42
36
03
Milk → Microwave (close)
60
56
40
50
04
Bottle → Cabinet (close)
78
76
70
78
05
Wine → Cabinet (close)
40
58
72
58
Appendix
Table 15: Per-task ASR on RoboCasa Tabletop / GR1. Success rates (%) over 50 episodes per task. Ours denotes RESETTLE; bold marks the better result within each model pair, including ties.
Task
Direct π0.5
Harness VLA
Base.
Ours
Base.
Ours
0
0
10
70
80
1
40
50
60
70
2
10
20
50
50
3
0
0
0
0
4
30
70
50
80
Appendix
Table 16: Per-task ASR on LIBERO-Pro Swap. Success rates (%) over 10 episodes per task. Ours denotes RESETTLE; bold marks the better result within each route, including ties.
Robotic manipulation poses fundamental challenges due to uncertainty, long-horizon execution, and compounding errors, which can easily destabilize execution and lead to task failure. Although recent vision-language-action (VLA) models exhibit strong generalization, they typically lack explicit mechanisms to assess execution stability and to recover when execution deviates from its nominal behavior. In this paper, we propose: (1) two complementary metrics to assess execution quality at runtime, and (2) an agentic reinforcement learning framework that learns to restore effective execution through high-level decision-making rather than directly learning low-level actions. In this framework, an agentic policy reasons over recent execution history and selects among a small set of execution modes to regulate the execution process. Under execution degradation, it triggers appropriate recovery mechanisms to restore the robot to previously visited nominal states, enabling the task to continue. We evaluate the proposed method on the LIBERO benchmark, achieving up to a 13.7% improvement in success rate under standard settings and up to a 39.2% improvement under disturbance settings, demonstrating substantially enhanced execution robustness.
Xiaopeng Zhang, Yueyang Weng, Qi Liu +2
School of Inteligence Science and Engineering, the Harbin Institute of Technology Shenzhen, 518055, China. · Faculty of Robot Science and Engineering, Northeastern University, Shenyang, 110819, China.
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100% success rates, seemingly suggesting that the models are ready for deployment in real world. However, near perfect performance on existing benchmarks can be misleading: success under ideal conditions does not imply real world robustness. Existing benchmarks primarily evaluate task completion from predefined initial states, while real world interactions inevitably involve failures such as failed grasps, collisions, and unintended object movements. A robot must therefore not only execute tasks successfully, but also recognize and recover from failures to continue the task. Yet this capability remains largely unmeasured, revealing a critical gap between benchmark performance and real world reliability. To address this gap, we introduce LIBERO-Recover Benchmark, a large scale benchmark for failure recovery in robotic manipulation. Built upon LIBERO, we collect real execution failures from SOTA embodied models and construct 1,000+ scenarios across four recovery levels: (1) Action Retry, (2) Action Adaptation, (3) Object State Recovery, and (4) Environmental Recovery. We evaluate four core capabilities: spatial understanding, object structure reasoning, interaction understanding, and topological reasoning. As the first large-scale benchmark for embodied failure recovery, LIBERO-Recover shifts evaluation from \emph{Can the robot succeed?''} to \emph{Can the robot recover after failure?''}, promoting robust and generalizable embodied agents. The project will be avaible in \textcolor{blue}{https://liulin815.github.io/LIBERO-Recovery/}.
Lin Liu, Zhicheng Bao, Lu Zhang +7
Beta Infinity · Beijing Jiaotong University · School of Information and Communication Engineering Dalian University of Technology +1
Vision-Language-Action (VLA) policies achieve strong performance in robotic manipulation but remain brittle once execution deviates from nominal trajectories. We propose CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution. Instead of generating corrective data from manually designed or random perturbations, CARE collects failed rollouts, models stage-conditioned post-failure deviations, and uses the resulting empirical distributions to synthesize representative failure states and corrective demonstrations. At inference time, CARE combines stage-wise planning with physically grounded 3D monitoring to trigger atomic adjustments or re-operations while preserving task progress. We further introduce the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies. Experiments across multiple VLA backbones, simulation benchmarks, and real-world dual-arm tasks show consistent improvements, with average task-success gains of 14.5 points in simulation and 15.9 points in the real world. Code, models, and data are available at https://github.com/xiaojunlan/care