The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.
Figures & tables
Figure 1: Comparison of policy-improvement paradigms. F4R forms an autonomous closed loop that recognizes deployment failures, reconstructs them in simulation, refines the policy, and redeploys it without human intervention. In contrast, targeted behavior cloning follows an open-loop pipeline that requires manual diagnosis and repeated real-world data recollection, resulting in substantial human effort and limited scalability.
Figure 2: Overview of the F4R architecture. (A) An initial policy is deployed in the real world to collect failure trajectories and task-relevant observations. (B) F4R reconstructs the failure scene in two stages. The resulting PBR meshes drive physical simulation, while the aligned Gaussians enable photorealistic rendering. (C) The policy first acquires corrective behaviors through failure-conditioned sim-real co-training and then improves them through targeted RL in the reconstructed environments. (D) The refined policy is redeployed and discovers new failure cases, closing the real-to-sim-to-real learning loop.
Method
In Distribution
ID Avg.
Out of Distribution
OOD Avg.
Pick Fruits
Place Cup on Coaster
Stack Bowls
Place Block in Drawer
Pick Fruits
Place Cup on Coaster
Stack Bowls
Place Block in Drawer
Base
70%
60%
50%
70%
62.5%
40%
35%
30%
0%
26.25%
Targeted BC
100%
95%
80%
90%
91.25%
85%
80%
60%
60%
71.25%
RLinf-Co
95%
85%
80%
90%
87.5%
90%
70%
70%
80%
77.5%
F4R (Ours)
100%
95%
85%
95%
93.75%
100%
95%
80%
85%
90%
Table 1: Success rates on four representative tasks under the original distribution and the deployment-derived failure distribution.
Figure 3: Sim–real behavioral consistency. Reconstructed environments for eight tabletop tasks and simulation versus real-world success rates for 24 policy checkpoints. Each marker represents one task and demonstration scale.
Figure 4: Multi-round closed-loop policy refinement. Real-world success rates of the base policy and the refined policies after three consecutive F4R cycles.
Table S1: Diagnosis taxonomy and structured output.
Parameter
Value
Definition
Scan stride s
3 frames
Change-estimation interval
Scan resolution
160×90
Grayscale proposal input
Smoothing S5
5 samples
Before and after MAD scaling
Peak threshold
qi,t>0
Positive standardized change
Maximum peaks
10
Highest-scoring motion peaks
Pre/post span
120/180 frames
Peak-window expansion
Table S2: Temporal-proposal parameters.
Figure S1: Temporal evidence board. The board contains 16 ordered time points in this example. Each panel pairs synchronized wrist and global observations. Adaptive sampling becomes denser around the candidate interaction while retaining interval-level context. Event boards contain up to 24 time points.
Figure S2: Full-rollout storyboards from the in-house dataset. Each board uniformly samples 16 synchronized wrist/external pairs. The original frame index and timestamp remain visible in every panel.
camera viewpoint, illumination, occluder placement
Table S4: Diagnosis-to-randomization mapping implemented by the diagnostic skill.
Script
Function
Primary output
detect_event_windows.py
Dual-view change scanning and window merging
event_windows.json/.csv
make_event_storyboards.py
Adaptive event-board construction
event JPEGs and manifest
make_state_comparisons.py
First-eight/last-eight state-board construction
state JPEGs and manifest
analyze_failure_v2.py
Window, state, focused, and aggregate VLM calls
cached and sample-level JSON
export_review_csv.py
Human-readable diagnostic audit
review_v2.csv
export_failure_conditioned_dr.py
Failure-aware reset and randomization export
failure_conditioned_dr_v2.json
Table S5: Diagnostic-skill scripts and generated artifacts.
Figure S3: Representative diagnosis cases. The first row shows three correctly diagnosed rollouts, while the second row shows two incorrect ViFailback cases. A diagnosis is correct only when both cause and failed subtask match the annotation. The error cases substitute a salient downstream consequence for an earlier causal failure.
Result
Sample
Prediction / annotation
Frame (pred./GT)
Evidence or error
Correct
Coke can
gripper state, subtask 2 / same
152 / 112
Can remains on table after gripper withdrawal
Correct
Spatula
gripper state, subtask 2 / same
120 / 75
Arm transports while spatula remains on table
Correct
Green cube
gripper state, subtask 2 / same
203 / 137
Transport continues after cube is lost
Error
Duck
task planning, subtask 1 / 6D pose, subtask 3
50 / 175
Similar object masks later pose error near cup
Error
Marker
6D pose, subtask 3 / task planning, subtask 1
130 / 63
Rim placement is mistaken for the first cause
Table S6: Qualitative ViFailback audit. Prediction/annotation pairs for the five visualized cases.
Outcome
Count
Percentage
Correct after one attempt
69/80
86.25%
Correct within two attempts
74/80
92.50%
Remaining: camera viewpoint
4/80
5.00%
Remaining: illumination
2/80
2.50%
Table S7: Autonomous failure diagnosis. Accuracy after one and two attempts on 80 trajectories across eight tasks.
Figure S4: Complete ViFailback results. Failure-cause accuracy and end-to-end inference speed of mainstream VLM configurations without benchmark-specific adaptation.
Figure S5: Manipulation-scene reconstruction. The scene is first reconstructed into a 3D Gaussian representation from captured videos. After pruning reconstruction outliers, the Gaussian primitives are aligned with the simulated robot using a scale-aware affine transformation. The aligned robot mesh is then used to assign link-level labels to the Gaussians, enabling the simulation-driven motion of Gaussian representations.
Setting
Value
Smartphone model
IPhone 17 Pro
Capture resolution / frame rate
1920 × 1080 / 30 FPS
Capture duration / frame count
≈ 5 min
3DGS optimization iterations
30000
3DGS optimization hyperparameters
Default
Scene-to-simulator registration method
FPFH → Sim3 RANSAC → Sim3 ICP
Table S8: Scene reconstruction settings.
Figure S6: Scene reconstruction quality. The top row shows images captured from the first-person perspective using a real-world RealSense D435i camera; the second row displays images rendered using Gaussian Splatting, with the SSIM between the sim-real images listed in the top-right corner; the third row shows the frame difference.
Figure S7: Object reconstruction quality. Rendering of object assets using in the experiments.
Parameter
Value
Input depth range
[0.05,2.0] m
Back-projection stride
1 px
Per-frame voxel size
0.003 m
Normal radius
3.0× voxel size
Normal max neighbors
30
Frame-pair window
3 neighboring frames
Table S9: Object reconstruction settings.
Stage
Human involvement
Frequency
Typical time
Workspace smartphone capture
One-time manual
New workspace
3–5 min
Scene reconstruction and alignment
Lightweight verification
New/changed workspace
∼ 20 min
Wrist RGB-D object observation
Automatic robot execution
New object/failure case
∼ 2 min
Object asset reconstruction
Automatic
New object
5–10 min
Corrective data generation
Automatic
Each refinement cycle
∼ 10 min
Real-world corrective demonstration
Not required
—
0 min
Table S10: Human-effort breakdown. Automation boundary of F4R.
Task / failure
Variable
Nominal source
Feasibility constraint
Pick Fruits
Fruit pose (x,y,ψ)
diagnosed scene
reachable; no overlap
Pick Fruits
Basket pose / visibility
diagnosed scene
target remains valid
Place Cup on Coaster
Cup pose
diagnosed scene
reachable
Place Cup on Coaster
Coaster pose / wrist visibility
diagnosed scene
valid placement area
Stack Bowls
Bowl poses / placement order
diagnosed scene
stable initial state
Place Block in Drawer
Block pose
diagnosed scene
reachable
Table S11: Failure-aware randomization ranges. Active variables and sampling distributions for the four refinement tasks.
Task
Historical seeds
Generated
Successful
Retained / generated
AnyGrasp used
Final ∣DsimF∣
Avg. horizon
Pick Fruits
24
500
438
82.0%
64 / 410
410
62.5
Place Cup on Coaster
22
560
462
76.8%
118 / 430
430
71.3
Stack Bowls
28
640
486
70.3%
156 / 450
450
83.7
Place Block in Drawer
26
620
421
61.3%
142 / 380
380
96.4
Table S12: Corrective-data generation statistics. Historical seeds denote real-world successful rollouts used to initialize MimicGen-style generation. Generated denotes all candidate simulated rollouts before filtering. Successful denotes rollouts satisfying the simulator success predicate. Final ∣DsimF∣ denotes the number of trajectories retained for supervised co-training.
Stage
Hyperparameter
Value
SFT
Base policy
π0.5
SFT
Trainable modules
LoRA
SFT
Optimizer
AdamW
SFT
Learning rate
2.5×10−5
SFT
GPUs
4 × NVIDIA H20
SFT
Distributed strategy
PyTorch FSDP
Table S13: SFT and PPO hyperparameters.
Figure S8: Real-robot platform. The dual-view perception system combines global task context with local manipulation observations.
Setting
Value
Robot / gripper
UR5e / Robotiq 2F-85
Cameras
2 × Intel RealSense D435i
Camera placement
Fixed third-person and wrist-mounted views
Camera capture
640×480 RGB at 30 Hz
Policy image resolution
224×224 RGB with aspect-ratio-preserving resize
Hand–eye calibration
Target-based calibration; extrinsics fixed across trials
Table S14: Robot, camera, and control settings. Hardware, observation, calibration, control, and safety configurations used in all real-world experiments.
Task
Language instruction
Objects
Success criterion
Horizon (steps)
Pick Fruits
Place the fruit in the basket.
fruit, basket
Fruit remains stably contained in the basket
1500
Place Cup on Coaster
Place the cup on the coaster.
cup, coaster
Cup remains stably supported by the coaster
1500
Hang Cup
Hang the cup on the rack.
cup, rack
Cup remains suspended from the rack after release
1500
Place Cup in Bowl
Place the cup in the bowl.
cup, bowl
Cup remains stably contained in the bowl
1500
Stack Blocks
Stack one block on the other.
two blocks
Upper block remains stably supported by the lower block
1500
Insert Cylinder into Board
Insert the cylinder into the board.
cylinder, insertion board
Cylinder remains inserted in the designated hole
1500
Table S15: Eight-task benchmark. Language instructions, objects, binary success criteria, episode horizons, and numbers of real-world demonstrations used to train the evaluated checkpoints.
Figure S9: Real–simulation paired views. The reconstructed environments preserve the task-relevant objects, spatial relations, and viewpoints across all eight tasks.
Setting
Translation range
Yaw range
Feasibility constraint
Trials per task
Shared init.
ID
x∼U(0.02,0.88) m, y∼U(−0.185,0.957) m, z=z0
ψ∼U(−π/6,π/6)
Reachable and collision-free
20
Yes
OOD
Same feasible workspace
ψ∈[−π/3,−π/6)ψ∈∪(π/6,π/3]
Reachable and collision-free
20
Yes
Table S16: ID and OOD evaluation distributions. Positions are expressed in the robot workspace frame. OOD cases use deployment-derived failure orientations outside the ID yaw range while satisfying the same reachability and collision constraints.
Method
Human corrective data
Scene setup
Object setup
Sim collection
Training
Reusable assets
Targeted BC
60 min
0
0
0
15h
—
RLinf-Co
0
shared 25–30 min
5–10 min
10 min
10-14h
Yes
F4R
0
shared 25–30 min
5–10 min
10 min
10-14h
Yes
Later F4R cycle, unchanged workspace/object
0
0
0
10 min
10-14h
Reused
Table S17: Budget and cost breakdown. Initial-cycle and amortized costs under the matched comparison.
Task
Collection time
Successful demos
Failed attempts
Mean length
Pick Fruits
1.0 h
48
2
570
Place Cup on Coaster
1.0 h
42
3
630
Stack Bowls
1.0 h
24
3
1,740
Place Block in Drawer
1.0 h
35
2
860
Table S18: Per-task Targeted BC data collection. Failed attempts are excluded from the training data. Trajectory length is measured as the number of transitions recorded at 30 Hz.
Task
Demonstra- tions
Real success (%)
Sim success (%)
∣Δ∣ (pp)
Trials (real/sim)
Pick Fruits
10
20
25
5
20/100
Pick Fruits
30
50
47
3
20/100
Pick Fruits
50
100
77
23
20/100
Put Cup on Coaster
10
20
17
3
20/100
Put Cup on Coaster
30
70
60
10
20/100
Put Cup on Coaster
50
95
76
19
20/100
Table S20: Complete sim–real consistency results. Raw success rates for all 24 checkpoints. The sim–real gap ∣Δ∣ is reported in percentage points (pp).
Figure S10: Evaluation on multi tasks.
Component
Failure mode
Consequence
Possible mitigation
Diagnosis
Subtle viewpoint or illumination change
Incorrect causal attribution
Longer temporal memory and explicit view normalization
Diagnosis
Multiple simultaneous failures
Ambiguous earliest cause
Causal multi-hypothesis diagnosis
Reconstruction
Transparent, reflective, or thin objects
Incomplete geometry/depth
Specialized sensing or material-aware reconstruction
Reconstruction
Ambiguous articulation
Incorrect joint axis/limits
Active articulation probing
Simulation
Friction/contact mismatch
Sim–real execution gap
Online parameter identification
Simulation
Deformable objects or dynamic scenes
Invalid rigid-scene assumption
Deformable/dynamic reconstruction
Table S22: Current limitations and possible mitigations.
Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many existing recovery methods rely on human intervention. In this paper, we propose Failure-Aware Retry (FAR), a framework that enables robots to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete the task autonomously. FAR combines Failure-Contrastive Preference Adaptation, which constructs preference learning data from failures to steer the policy away from previously unsuccessful behaviors, with lightweight action perturbations during retries to encourage local exploration. We further incorporate successful recovery trajectories into a training loop for continual policy improvement. Experiments in both simulation and real-world manipulation tasks show that FAR substantially improves success rates and robustness, with average gains of 17.6% over the standard diffusion policy in simulation and 11.7% in the real world. In addition, FAR significantly improves data efficiency under both reset and timestep budgets during continual policy improvement by exploiting informative failure cases.
Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose FailPatch, a failure-driven residual patching framework that decouples action supervision from execution-reliability supervision. Successful demonstrations ground how the policy should act, while deployment trajectories indicate when its behavior becomes unreliable. We further observe that action hidden representations exhibit clear linear separability between reliable and failure-associated states while directly conditioning action generation. Building on these insights, FailPatch introduces a Null-gated Residual Expert Bank into the action hidden space of a frozen VLA policy. A unified Preserve--Redirect--Trust objective retains the original policy in reliable states, selects residual experts in failure-associated states and redirects representations from failure regions toward success-associated regions under bounded intervention. With only 0.52% trainable parameters, FailPatch improves success rates by 11.0 percentage points on four long-horizon RoboTwin tasks under clean evaluation, 9.5 percentage points under clean-to-random generalization, and 16.7 percentage points over the baseline across three real-world tasks. Project and code: https://github.com/yupeng-2003/FailPatch.
Peng Yu, Jiacheng Wang, Ziheng Zhang +6
Xi’an Jiaotong University · Dexmal · Nanjing University
Vision-language-action (VLA) models provide a promising paradigm for scalable robotic manipulation, yet their reliance on success-only behavioral cloning leaves them brittle; lacking corrective training signals, minor execution errors rapidly compound into unrecoverable, out-of-distribution failures. To address this limitation, we propose Adaptive Failure-Informed Learning (AFIL), an end-to-end framework that leverages failure trajectories as adaptive negative guidance for diffusion- and flow-based VLA policies. AFIL uses a pretrained VLA to generate failure rollouts online, avoiding the need for handcrafted failure-mode design or human-in-the-loop recovery. It then jointly trains Dual Action Generators (DAGs) for successful and failed behaviors while sharing a common vision-language backbone, enabling efficient failure-aware policy learning with limited parameter overhead. During sampling, the failure generator adaptively steers action generation away from failure-prone regions and toward more reliable success modes, with guidance strength determined by the per-diffusion-step distance between success and failure distributions. Experiments across in-domain and out-of-domain robotic manipulation tasks, covering both short- and long-horizon settings, show that AFIL consistently improves task success rates and robustness over existing VLA baselines, demonstrating its effectiveness, efficiency, and generality.
Meng Zheng, Samhita Marri, Anwesa Choudhuri +6
United Imaging Intelligence, Boston, MA, USA · University of Illinois Urbana-Champaign, Urbana, IL, USA