RESETTLE: Robotic Recovery through Disagreement-Triggered Retrieval and Efficient Corrective Control
Organizations: The Hong Kong University of Science and Technology · The Chinese University of Hong Kong · SmartMore · Harbin Institute of Technology, Shenzhen · Shenzhen Loop Area Institute
Abstract
Reliable robotic manipulation requires timely intervention to correct emerging deviations and restore progress after execution errors. However, recovery methods based on repeated vision-language reasoning or iterative online optimization can incur substantial latency, delaying intervention. To address these challenges, we introduce RESETTLE(Robotic rEcovery through diSagrEement-Triggered reTrievaL and Efficient Corrective Control), a model-agnostic framework that provides computationally efficient recovery at the action-execution interface of frozen robot policies. RESETTLE triggers recovery when two action proposals independently sampled under identical conditioning persistently disagree. It retrieves a same-task demonstration reference using an adapted V-JEPA encoder and combines a state-servo prior with a guarded visual residual to execute one corrective action without online trajectory optimization or additional vision-language reasoning, then returns control to the base policy. Across six base policies in simulation, RESETTLE achieves up to 8.70%, 6.28%, and 6.83% absolute success-rate gains on LIBERO-Plus, Meta-World, and RoboCasa Tabletop, respectively, with further improvements on four real-world tasks using two policies. In QwenPI-based comparisons, its monitoring-and-recovery computation latency is 74.04%--93.57% lower than VoLoAgent's monitoring-and-planning latency for grasp and place tool calls. It also raises Harness VLA's LIBERO-Pro Swap success from 42% to 50%, demonstrating compatibility with high-level agentic planning. Code available at: https://github.com/JIA-Lab-research/RESETTLE
Figures & tables
| Model | Cam. | Robot | Lang. | Light | BG | Noise | Layout | Average |
|---|---|---|---|---|---|---|---|---|
| OpenVLA | 0.8 | 3.5 | 23.0 | 8.1 | 34.8 | 15.2 | 28.5 | 15.6 |
| OpenVLA-OFT | 56.4 | 31.9 | 79.5 | 88.7 | 93.3 | 75.8 | 74.2 | 69.6 |
| NORA | 2.2 | 37.0 | 65.1 | 45.7 | 58.6 | 12.8 | 62.1 | 39.0 |
| WorldVLA | 0.1 | 27.9 | 41.6 | 43.7 | 17.1 | 10.9 | 38.0 | 25.0 |
| UniVLA | 1.8 | 46.2 | 69.6 | 69.0 | 81.0 | 21.2 | 31.9 | 42.9 |
| 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.8 | 53.6 |
| Model | Easy (28) | Medium (11) | Hard (6) | Very Hard (5) | Average |
|---|---|---|---|---|---|
| QwenPI / Base. | 80.36 | 57.27 | 45.00 | 60.00 | 60.66 |
| QwenPI / Ours | 83.57 | 62.73 | 51.67 | 66.00 | 65.99 |
| SmolVLA / Base. | 79.64 | 41.82 | 83.33 | 32.00 | 59.20 |
| SmolVLA / Ours | 83.57 | 46.36 | 90.00 | 42.00 | 65.48 |
| Model | Press-button | Cube-up | Carrot-in-pot | Use Spoon | Avg. |
|---|---|---|---|---|---|
| SR | SR | SR | SC | ||
| QwenPI / Base. | 90 | 55 | 85 | 36.7 | 66.67 |
| QwenPI / Ours | 100 | 70 | 95 | 50 | 78.75 |
| VLAct / Base. | 90 | 85 | 90 | 80 | 86.25 |
| VLAct / Ours | 100 | 90 | 95 | 86.7 | 92.92 |
| Method | Monitor | Recovery | Total | Relative latency |
|---|---|---|---|---|
| RESETTLE (Ours) | 0.0 | 83.6 | 83.6 | |
| Agentic Robot | 107.0 | 109.0 | 216.0 | |
| VoLoAgent | 107.0 | 215.0–1193.0 | 322.0–1300.0 | – |
| LPB † | 56.2 | 423.5 | 479.7 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Module | Total parameters | Trainable parameters |
|---|---|---|
| Context encoder | 1,012,173,952 | 1,012,173,952 |
| EMA target encoder | 1,012,173,952 | 0 |
| Action-conditioned predictor | 305,220,992 | 305,213,824 |
| Guarded residual-action head | 6,839,226 | 6,839,226 |
| All instantiated modules | 2,336,408,122 | 1,324,227,002 |
| Deployed subset | 1,019,013,178 | – |
| Setting | LIBERO | RoboCasa | Meta-World |
|---|---|---|---|
| State/action dim. | 7 | 29 | 4 |
| Frames per clip | 8 | 8 | 8 |
| Input resolution | |||
| Per-GPU batch | 16 | 16 | 16 |
| Global batch | 128 | 128 | 128 |
| Training budget | 20k updates | 150k updates | 20 epochs |
| Perturbation | V-JEPA Feature Distance | Action Disagreement (Ours) | ||||||
|---|---|---|---|---|---|---|---|---|
| Failed | Successful | Failed | Successful | |||||
| Rate (%) | Count | Rate (%) | Count | Rate (%) | Count | Rate (%) | Count | |
| Camera | 26.52 | 179.3 | 10.00 | 66.9 | 93.10 | 16.6 | 29.47 | 2.4 |
| Robot | 27.98 | 104.6 | 17.44 | 28.2 | 95.40 | 16.9 | 25.40 | 2.6 |
| Language | 15.48 | 70.5 | 17.89 | 24.8 | 77.89 | 9.6 | 9.22 | 1.9 |
| Light | 21.74 | 10.2 | 9.74 | 59.4 | 91.67 | 12.6 | 12.34 | 1.8 |
| Setting | Failed | Successful | ||
|---|---|---|---|---|
| Rate (%) | Count | Rate (%) | Count | |
| FastWAM / LIBERO-Plus | ||||
| Camera | 42.01 | 2.0 | 12.99 | 1.5 |
| Robot | 53.36 | 2.4 | 11.19 | 1.4 |
| Language | 68.99 | 3.5 | 8.91 | 1.4 |
| Light | 64.83 | 3.0 | 10.71 | 1.2 |
| Perturbation | Mean takeovers per trajectory | Intervened trajectories (%) | Recovery actions (%) | |
|---|---|---|---|---|
| All | Intervened | |||
| Camera | 3.90 | 8.80 | 44.28 | 1.62 |
| Robot | 25.32 | 29.73 | 85.16 | 12.59 |
| Language | 0.07 | 2.78 | 2.34 | 0.04 |
| Light | 0.07 | 2.58 | 2.71 | 0.05 |
| Background | 0.29 | 2.70 | 10.78 | 0.19 |
| Setting | Cam. | Robot | Lang. | Light | BG | Noise | Layout | Total |
|---|---|---|---|---|---|---|---|---|
| Baseline | 55.7 | 61.0 | 94.5 | 98.0 | 97.6 | 86.2 | 79.1 | 80.2 |
| P85 | 58.7 | 84.2 | 94.7 | 97.5 | 97.7 | 79.5 | 78.8 | 83.1 |
| P90 | 59.2 | 83.8 | 94.3 | 98.0 | 97.9 | 83.3 | 79.3 | 83.9 |
| P95 | 59.9 | 82.4 | 94.7 | 97.8 | 98.6 | 86.9 | 80.1 | 84.6 |
| P99 | 57.3 | 76.2 | 94.9 | 97.5 | 98.5 | 89.6 | 80.1 | 83.6 |
| ID | Task | QwenPI | SmolVLA | ||
|---|---|---|---|---|---|
| Base. | Ours | Base. | Ours | ||
| 00 | nut-assembly | 40 | 40 | 90 | 90 |
| 01 | basketball | 30 | 30 | 10 | 10 |
| 02 | bin-picking | 40 | 40 | 10 | 20 |
| 03 | box-close | 30 | 70 | 40 | 50 |
| 04 | button-press-topdown | 100 | 100 | 100 | 100 |
| ID | Task | QwenPI | SmolVLA | ||
|---|---|---|---|---|---|
| Base. | Ours | Base. | Ours | ||
| 25 | handle-pull-side | 10 | 20 | 20 | 30 |
| 26 | handle-pull | 10 | 30 | 100 | 100 |
| 27 | lever-pull | 40 | 40 | 0 | 10 |
| 28 | pick-place-wall | 70 | 70 | 40 | 20 |
| 29 | pick-out-of-hole | 0 | 30 | 40 | 30 |
| ID | Task | QwenGR00T | LDA | ||
| Base. | Ours | Base. | Ours | ||
| 01 | Cup Drawer (close) | 36 | 34 | 50 | 42 |
| 02 | Potato Microwave (close) | 30 | 42 | 42 | 36 |
| 03 | Milk Microwave (close) | 60 | 56 | 40 | 50 |
| 04 | Bottle Cabinet (close) | 78 | 76 | 70 | 78 |
| 05 | Wine Cabinet (close) | 40 | 58 | 72 | 58 |
| Task | Direct | Harness VLA | ||
|---|---|---|---|---|
| Base. | Ours | Base. | Ours | |
| 0 | 0 | 10 | 70 | 80 |
| 1 | 40 | 50 | 60 | 70 |
| 2 | 10 | 20 | 50 | 50 |
| 3 | 0 | 0 | 0 | 0 |
| 4 | 30 | 70 | 50 | 80 |