FailPatch: Failure Residual Patching for Vision-Language-Action Models
Organizations: Xi’an Jiaotong University · Dexmal · Nanjing University
Abstract
Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose FailPatch, a failure-driven residual patching framework that decouples action supervision from execution-reliability supervision. Successful demonstrations ground how the policy should act, while deployment trajectories indicate when its behavior becomes unreliable. We further observe that action hidden representations exhibit clear linear separability between reliable and failure-associated states while directly conditioning action generation. Building on these insights, FailPatch introduces a Null-gated Residual Expert Bank into the action hidden space of a frozen VLA policy. A unified Preserve--Redirect--Trust objective retains the original policy in reliable states, selects residual experts in failure-associated states and redirects representations from failure regions toward success-associated regions under bounded intervention. With only 0.52% trainable parameters, FailPatch improves success rates by 11.0 percentage points on four long-horizon RoboTwin tasks under clean evaluation, 9.5 percentage points under clean-to-random generalization, and 16.7 percentage points over the baseline across three real-world tasks. Project and code: https://github.com/yupeng-2003/FailPatch.
Figures & tables
| Method | Blocks Ranking | Stack Bowls | Put Object Cabinet | Place Bread Basket | Avg. |
| Without rollout trajectories | |||||
| 50-Task Base | 38 | 68 | 40 | 52 | 49.50 |
| LoRA SFT | 44 | 68 | 36 | 54 | 50.50 |
| Full SFT | 44 | 70 | 42 | 55 | 52.75 |
| Demo-Only REB | 42 | 64 | 36 | 55 | 49.25 |
| With rollout trajectories | |||||
| Method | Blocks Ranking | Stack Bowls | Object Cabinet | Bread Basket | Avg. |
|---|---|---|---|---|---|
| 50-Task Base | 12 | 39 | 25 | 36 | 28.00 |
| LoRA SFT | 13 | 33 | 20 | 33 | 24.75 |
| Full SFT | 23 | 45 | 29 | 34 | 32.75 |
| Demo-Only REB | 20 | 39 | 31 | 32 | 30.50 |
| Outcome SFT (Ep.) | 10 | 29 | 28 | 38 | 26.25 |
| Outcome SFT (Ch.) | 16 | 35 | 32 | 32 | 28.75 |
| Method | Blocks | Stack | Fold | Avg. |
|---|---|---|---|---|
| Ranking | Bowls | Dishcloth | ||
| Three-Task Base | 36.7 | 53.3 | 33.3 | 41.1 |
| LoRA SFT | 46.7 | 50.0 | 40.0 | 45.6 |
| Demo-Only REB | 43.3 | 50.0 | 36.7 | 43.3 |
| FailPatch | 56.7 | 66.7 | 50.0 | 57.8 |
| Variant | Avg. Success (%) |
|---|---|
| Objective components | |
| Demo-Only REB | 49.25 |
| + Preserve | 52.50 |
| + Redirect | 56.50 |
| + Trust | 50.25 |
| + Preserve + Redirect | 59.00 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Algorithm 1: FailPatch offline preparation, joint training, and inference | |
|---|---|
| Input: frozen policy and action projection ; demonstrations ; deployment rollouts ; experts ; router . | |
| 1 | Given the terminal outcome , annotate each rollout to obtain failure onset for failed rollouts and success reliability for successful rollouts. |
| 2 | Cache for all deployment chunks. Form from reliable-success chunks and failed-rollout pre-onset chunks; form from onset/post-onset chunks. |
| 3 | For each task , construct from normalized chunks of reliable successful rollouts only. |
| 4 | For : |
| 5 | Sample task-balanced chunks for the demonstration, Preserve, and Redirect branches according to the implementation, without trajectory-duration reweighting. |
| Configuration | Value |
| Number of residual experts | 8 |
| Nearest neighbors | 20 |
| Demo routing weight | 1.0 |
| Preserve weight | 0.2 |
| Redirect weight | 0.2 |
| Trust weight | 0.05 |
| Setting | Task | Successful rollouts | ||
|---|---|---|---|---|
| Simulation | Blocks Ranking | 50 | 44 | 6 |
| Simulation | Stack Bowls | 50 | 46 | 4 |
| Simulation | Put Object Cabinet | 50 | 43 | 7 |
| Simulation | Place Bread Basket | 50 | 45 | 5 |
| Real world | Blocks Ranking | 50 | 44 | 6 |
| Real world | Stack Bowls | 50 | 45 | 5 |
| Task | Exact | MAE | Bias | |||
|---|---|---|---|---|---|---|
| Blocks Ranking | 50 | 28 | 40 | 47 | 0.72 | |
| Stack Bowls | 50 | 30 | 42 | 48 | 0.68 | |
| Put Object Cabinet | 50 | 26 | 39 | 48 | 0.80 | |
| Place Bread Basket | 50 | 28 | 39 | 47 | 0.76 | |
| Overall | 200 | 112 | 160 | 190 | 0.74 |
| Task | Max steps | Max chunks | Success predicate |
|---|---|---|---|
| Blocks Ranking | 1200 | 24 | For the two adjacent block pairs (red, green) and (green, blue), m and m; the centers satisfy ; and both grippers are open. |
| Stack Bowls | 1200 | 24 | After sorting the bowls by height, every pair has planar center distance below 0.04 m; their heights are within 0.02 m of ; and both grippers are open. |
| Put Object Cabinet | 700 | 14 | The planar distance from the object to the drawer functional point is below 0.05 m; its relative height satisfies m; and the executing gripper is open. |
| Place Bread Basket | 700 | 14 | Every bread object’s planar distance to the basket center is below 0.05 m, its height exceeds , and both grippers are open. |
| Task | Randomized task instance |
|---|---|
| Blocks Ranking | Each block is initialized with m, m, m, and rotation about the vertical axis up to 0.75 rad. The block side length is sampled uniformly from m. Target ranges are , , and m for red, green, and blue, respectively, with common target m. |
| Stack Bowls | Bowl centers are initialized with m and m. Bowl identities are assigned according to their initial ordering, and the target stack center is m. |
| Put Object Cabinet | The cabinet has m and m. The manipulated object and model instance are sampled from ten object categories. The object has m, m, and yaw up to ; when the initially sampled m, the implementation widens the admissible range to m. |
| Place Bread Basket | The basket center is fixed at m, while its yaw is sampled from and its model ID from . Each instance contains one or two bread objects with m, m, yaw up to , and model ID sampled from . |
| Attribute | Clean | Random |
|---|---|---|
| Background | Default wall and table textures; clean-background probability 1.0. | Unseen evaluation texture library; clean-background probability 0.02. |
| Table clutter | Disabled. | Enabled, with up to approximately ten additional distractor objects. |
| Table height | No perturbation. | Downward height perturbation of up to 0.03 m. |
| Lighting | Fixed. | Randomized; extreme lighting is sampled with probability 0.02. |
| Head-camera distance | No distance perturbation. | No distance perturbation. |
| Camera observations | One D435 head view and two wrist views. | Identical camera configuration. |
| Method | Trainable params | Fraction | Time | Offline rollouts/task | Fresh rollouts |
|---|---|---|---|---|---|
| LoRA SFT | 50.0M | 1.49% | 8.3 h | 0 | 0 |
| Full SFT | 3353.4M | 100% | 11.2 h | 0 | 0 |
| Demo-Only REB | 17.3M | 0.52% | 3.1 h | 0 | 0 |
| Outcome-Prompt SFT (Episode) | 17.3M | 0.52% | 5.5 h | 100 | 0 |
| Outcome-Prompt SFT (Chunk) | 17.3M | 0.52% | 4.7 h | 100 | 0 |
| PPO | 694M | 20.7% | 20.8 h | 0 | 12,800 |
| Prototype variant | Blocks | Stack | Cabinet | Bread | Average |
|---|---|---|---|---|---|
| Default: , median | 54 | 77 | 49 | 62 | 60.50 |
| 48 | 71 | 44 | 58 | 55.25 | |
| 53 | 76 | 48 | 61 | 59.50 | |
| 52 | 75 | 47 | 61 | 58.75 | |
| Mean aggregation | 52 | 75 | 47 | 60 | 58.50 |
| Unnormalized features | 51 | 74 | 46 | 60 | 57.75 |
| Hyperparameter | Evaluated values | Default | Average success (%) |
|---|---|---|---|
| Experts | 4, 8, 16 | 8 | 54.25, 60.50 , 58.00 |
| Preserve weight | 0.1, 0.2, 0.4 | 0.2 | 57.50, 60.50 , 58.25 |
| Redirect weight | 0.1, 0.2, 0.4 | 0.2 | 56.00, 60.50 , 59.25 |
| Trust weight | 0.02, 0.05, 0.10 | 0.05 | 58.75, 60.50 , 57.00 |
| Demo temperature | annealed, 0.2, 1.0 | annealed | 60.50 , 59.00, 55.50 |
| Failure temperature | 0.1, 0.2, 0.5 | 0.2 | 59.25, 60.50 , 57.75 |
| Variant | Blocks | Stack | Cabinet | Bread | Average |
|---|---|---|---|---|---|
| Demo-Only REB | 42 | 64 | 36 | 55 | 49.25 |
| Preserve | 50 | 63 | 38 | 59 | 52.50 |
| Redirect | 52 | 68 | 42 | 64 | 56.50 |
| Trust | 44 | 62 | 37 | 58 | 50.25 |
| Preserve Redirect | 54 | 74 | 48 | 60 | 59.00 |
| Parameter-Matched Single Expert | 50 | 71 | 44 | 58 | 55.75 |
| Task | Language instruction | Success criterion |
|---|---|---|
| Blocks Ranking | “Position red block, green block, and blue block in a left-to-right sequence, forming a row.” | The three blocks form one row in red–green–blue order from left to right. |
| Stack Bowls | “Stack the three bowls on the top of each other.” | The task ends successfully once all three bowls have been stacked; no additional fixed-duration stability test is imposed. |
| Fold Dishcloth | “Place the dishcloth in the center of the table, then fold it in half twice.” | The dishcloth is first placed at the table center and then successfully folded in half twice. |
| Feature | Balanced accuracy | ROC-AUC |
|---|---|---|
| Vision | 0.73 | 0.79 |
| Action hidden | 0.74 | 0.83 |
| Predicted action | 0.62 | 0.69 |
| Shuffled labels | 0.51 | 0.52 |