F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement
Organizations: Nanyang Technological University · Xi’an Jiaotong University · Dexmal
Abstract
The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.
Figures & tables
| Method | In Distribution | ID Avg. | Out of Distribution | OOD Avg. | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Pick Fruits | Place Cup on Coaster | Stack Bowls | Place Block in Drawer | Pick Fruits | Place Cup on Coaster | Stack Bowls | Place Block in Drawer | |||
| Base | 70% | 60% | 50% | 70% | 62.5% | 40% | 35% | 30% | 0% | 26.25% |
| Targeted BC | 100% | 95% | 80% | 90% | 91.25% | 85% | 80% | 60% | 60% | 71.25% |
| RLinf-Co | 95% | 85% | 80% | 90% | 87.5% | 90% | 70% | 70% | 80% | 77.5% |
| F4R (Ours) | 100% | 95% | 85% | 95% | 93.75% | 100% | 95% | 80% | 85% | 90% |
| Field | Meaning | Allowed values / example |
|---|---|---|
| first_failure_frame | Earliest visible anomaly | printed source-frame index |
| first_failure_window | Supporting evidence board | sample_w02 / final_state |
| target_object | Object manipulated at first failure | task-specific object name |
| operation | First failed stage | grasp, transport, place, release, uncertain |
| failure_type | Visible execution outcome | place/grasp failure, outside, slip, tip, collision, occlusion |
| matched_labels | Failure-relevant scene factors | pose, stiffness, friction, lighting, background, view, clutter |
| Parameter | Value | Definition |
|---|---|---|
| Scan stride | 3 frames | Change-estimation interval |
| Scan resolution | Grayscale proposal input | |
| Smoothing | 5 samples | Before and after MAD scaling |
| Peak threshold | Positive standardized change | |
| Maximum peaks | 10 | Highest-scoring motion peaks |
| Pre/post span | 120/180 frames | Peak-window expansion |
| Call | Core instruction | Required JSON fields |
|---|---|---|
| Event window | Identify visible accident or goal-level failure | window, evidence flag, type, frame range, object, operation, labels, evidence, confidence |
| State board | Compare initial and terminal states; do not infer motion | final-state failure, start objects, inside/outside objects, labels, evidence, confidence |
| Focused re-query | Find earliest anomaly in earliest positive board | first frame, target object, operation, type, labels, evidence, confidence |
| Aggregation | Verify the first cause using all preceding JSON | first failed subtask, critical events, labels, evidence, sample confidence |
| Diagnosed condition | Exported randomization targets |
|---|---|
| Object not placed in bowl | bowl position, opening orientation, object initial pose |
| Object slips during grasp | friction, object mass, contact parameters, gripper closing speed |
| Collision with nearby object | clutter density, obstacle position |
| Wrong object grasped | similar objects, color, texture, occlusion |
| Insertion failure | hole-position offset, peg angle, clearance, friction |
| Target occluded | camera viewpoint, illumination, occluder placement |
| Script | Function | Primary output |
|---|---|---|
| detect_event_windows.py | Dual-view change scanning and window merging | event_windows.json/.csv |
| make_event_storyboards.py | Adaptive event-board construction | event JPEGs and manifest |
| make_state_comparisons.py | First-eight/last-eight state-board construction | state JPEGs and manifest |
| analyze_failure_v2.py | Window, state, focused, and aggregate VLM calls | cached and sample-level JSON |
| export_review_csv.py | Human-readable diagnostic audit | review_v2.csv |
| export_failure_conditioned_dr.py | Failure-aware reset and randomization export | failure_conditioned_dr_v2.json |
| Result | Sample | Prediction / annotation | Frame (pred./GT) | Evidence or error |
|---|---|---|---|---|
| Correct | Coke can | gripper state, subtask 2 / same | 152 / 112 | Can remains on table after gripper withdrawal |
| Correct | Spatula | gripper state, subtask 2 / same | 120 / 75 | Arm transports while spatula remains on table |
| Correct | Green cube | gripper state, subtask 2 / same | 203 / 137 | Transport continues after cube is lost |
| Error | Duck | task planning, subtask 1 / 6D pose, subtask 3 | 50 / 175 | Similar object masks later pose error near cup |
| Error | Marker | 6D pose, subtask 3 / task planning, subtask 1 | 130 / 63 | Rim placement is mistaken for the first cause |
| Outcome | Count | Percentage |
|---|---|---|
| Correct after one attempt | 69/80 | 86.25% |
| Correct within two attempts | 74/80 | 92.50% |
| Remaining: camera viewpoint | 4/80 | 5.00% |
| Remaining: illumination | 2/80 | 2.50% |
| Setting | Value |
|---|---|
| Smartphone model | IPhone 17 Pro |
| Capture resolution / frame rate | 1920 1080 / 30 FPS |
| Capture duration / frame count | 5 min |
| 3DGS optimization iterations | 30000 |
| 3DGS optimization hyperparameters | Default |
| Scene-to-simulator registration method | FPFH Sim3 RANSAC Sim3 ICP |
| Parameter | Value |
|---|---|
| Input depth range | m |
| Back-projection stride | px |
| Per-frame voxel size | m |
| Normal radius | voxel size |
| Normal max neighbors | |
| Frame-pair window | neighboring frames |
| Stage | Human involvement | Frequency | Typical time |
|---|---|---|---|
| Workspace smartphone capture | One-time manual | New workspace | 3–5 min |
| Scene reconstruction and alignment | Lightweight verification | New/changed workspace | 20 min |
| Wrist RGB-D object observation | Automatic robot execution | New object/failure case | 2 min |
| Object asset reconstruction | Automatic | New object | 5–10 min |
| Corrective data generation | Automatic | Each refinement cycle | 10 min |
| Real-world corrective demonstration | Not required | — | 0 min |
| Task / failure | Variable | Nominal source | Feasibility constraint |
|---|---|---|---|
| Pick Fruits | Fruit pose | diagnosed scene | reachable; no overlap |
| Pick Fruits | Basket pose / visibility | diagnosed scene | target remains valid |
| Place Cup on Coaster | Cup pose | diagnosed scene | reachable |
| Place Cup on Coaster | Coaster pose / wrist visibility | diagnosed scene | valid placement area |
| Stack Bowls | Bowl poses / placement order | diagnosed scene | stable initial state |
| Place Block in Drawer | Block pose | diagnosed scene | reachable |
| Task | Historical seeds | Generated | Successful | Retained / generated | AnyGrasp used | Final | Avg. horizon |
|---|---|---|---|---|---|---|---|
| Pick Fruits | 24 | 500 | 438 | 82.0% | 64 / 410 | 410 | 62.5 |
| Place Cup on Coaster | 22 | 560 | 462 | 76.8% | 118 / 430 | 430 | 71.3 |
| Stack Bowls | 28 | 640 | 486 | 70.3% | 156 / 450 | 450 | 83.7 |
| Place Block in Drawer | 26 | 620 | 421 | 61.3% | 142 / 380 | 380 | 96.4 |
| Stage | Hyperparameter | Value |
|---|---|---|
| SFT | Base policy | |
| SFT | Trainable modules | LoRA |
| SFT | Optimizer | AdamW |
| SFT | Learning rate | |
| SFT | GPUs | 4 NVIDIA H20 |
| SFT | Distributed strategy | PyTorch FSDP |
| Setting | Value |
|---|---|
| Robot / gripper | UR5e / Robotiq 2F-85 |
| Cameras | 2 Intel RealSense D435i |
| Camera placement | Fixed third-person and wrist-mounted views |
| Camera capture | RGB at 30 Hz |
| Policy image resolution | RGB with aspect-ratio-preserving resize |
| Hand–eye calibration | Target-based calibration; extrinsics fixed across trials |
| Task | Language instruction | Objects | Success criterion | Horizon (steps) |
|---|---|---|---|---|
| Pick Fruits | Place the fruit in the basket. | fruit, basket | Fruit remains stably contained in the basket | 1500 |
| Place Cup on Coaster | Place the cup on the coaster. | cup, coaster | Cup remains stably supported by the coaster | 1500 |
| Hang Cup | Hang the cup on the rack. | cup, rack | Cup remains suspended from the rack after release | 1500 |
| Place Cup in Bowl | Place the cup in the bowl. | cup, bowl | Cup remains stably contained in the bowl | 1500 |
| Stack Blocks | Stack one block on the other. | two blocks | Upper block remains stably supported by the lower block | 1500 |
| Insert Cylinder into Board | Insert the cylinder into the board. | cylinder, insertion board | Cylinder remains inserted in the designated hole | 1500 |
| Setting | Translation range | Yaw range | Feasibility constraint | Trials per task | Shared init. |
|---|---|---|---|---|---|
| ID | m, m, | Reachable and collision-free | 20 | Yes | |
| OOD | Same feasible workspace | Reachable and collision-free | 20 | Yes |
| Method | Human corrective data | Scene setup | Object setup | Sim collection | Training | Reusable assets |
|---|---|---|---|---|---|---|
| Targeted BC | 60 min | 0 | 0 | 0 | 15h | — |
| RLinf-Co | 0 | shared 25–30 min | 5–10 min | 10 min | 10-14h | Yes |
| F4R | 0 | shared 25–30 min | 5–10 min | 10 min | 10-14h | Yes |
| Later F4R cycle, unchanged workspace/object | 0 | 0 | 0 | 10 min | 10-14h | Reused |
| Task | Collection time | Successful demos | Failed attempts | Mean length |
|---|---|---|---|---|
| Pick Fruits | 1.0 h | 48 | 2 | 570 |
| Place Cup on Coaster | 1.0 h | 42 | 3 | 630 |
| Stack Bowls | 1.0 h | 24 | 3 | 1,740 |
| Place Block in Drawer | 1.0 h | 35 | 2 | 860 |
| Task | Demonstra- tions | Real success (%) | Sim success (%) | (pp) | Trials (real/sim) |
|---|---|---|---|---|---|
| Pick Fruits | 10 | 20 | 25 | 5 | 20/100 |
| Pick Fruits | 30 | 50 | 47 | 3 | 20/100 |
| Pick Fruits | 50 | 100 | 77 | 23 | 20/100 |
| Put Cup on Coaster | 10 | 20 | 17 | 3 | 20/100 |
| Put Cup on Coaster | 30 | 70 | 60 | 10 | 20/100 |
| Put Cup on Coaster | 50 | 95 | 76 | 19 | 20/100 |
| Component | Failure mode | Consequence | Possible mitigation |
|---|---|---|---|
| Diagnosis | Subtle viewpoint or illumination change | Incorrect causal attribution | Longer temporal memory and explicit view normalization |
| Diagnosis | Multiple simultaneous failures | Ambiguous earliest cause | Causal multi-hypothesis diagnosis |
| Reconstruction | Transparent, reflective, or thin objects | Incomplete geometry/depth | Specialized sensing or material-aware reconstruction |
| Reconstruction | Ambiguous articulation | Incorrect joint axis/limits | Active articulation probing |
| Simulation | Friction/contact mismatch | Sim–real execution gap | Online parameter identification |
| Simulation | Deformable objects or dynamic scenes | Invalid rigid-scene assumption | Deformable/dynamic reconstruction |