Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models
Organizations: University of Notre Dame
Abstract
Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress. To address this challenge, we introduce FailBank, a four-stage self-evolving framework that converts runtime feedback into persistent policy improvement. During collection, a fixed CBF-based safety module serves as an observe-only teacher, producing counterfactual corrections while the policy remains in control. Outcome-aware admission then converts useful proposals into corrective targets and retains successful uncorrected actions as quiet anchors for guarded LoRA updates. We evaluate FailBank on the VLA-Arena benchmark across two difficulty levels and two VLA backbones. Compared with the base policies, FailBank improves the joint success-cost operating point. Across the two backbones, FailBank improves task success rate by 8.5 and 6.9 percentage points, while reducing policy-induced cumulative cost by 35.6% and 23.8%, respectively. Compared with runtime shielding, FailBank raises task success rate by 25.4 and 9.5 percentage points, while maintaining comparable policy-induced cumulative cost. These results show that runtime feedback can serve as persistent policy supervision rather than only as a temporary action constraint.
Figures & tables
| Base | AEGIS | FailBank | Base | AEGIS | FailBank | ||||||||||||||
| Level | Task | SR | CC | SR | CC | SR | CC | SR | CC | SR | CC | SR | CC | ||||||
| 1 | Apple | 90.0 | 9.12 | 0.12 | 82.0 | 8.20 | 0.00 | 90.3 | 9.21 | 0.21 | 54.0 | 5.40 | 0.00 | 82.0 | 8.20 | 0.00 | 59.0 | 5.90 | 0.00 |
| Lemon | 83.0 | 44.07 | 35.87 | 78.0 | 7.81 | 0.01 | 81.0 | 25.98 | 17.88 | 66.0 | 6.60 | 0.00 | 80.0 | 8.00 | 0.00 | 84.0 | 8.40 | 0.00 | |
| Mango | 77.0 | 74.08 | 66.38 | 68.0 | 21.08 | 14.28 | 93.7 | 23.12 | 13.76 | 92.0 | 11.34 | 2.14 | 82.0 | 9.40 | 1.20 | 93.0 | 12.39 | 3.09 | |
| Onion | 72.0 | 33.88 | 27.18 | 44.0 | 9.12 | 4.71 | 90.0 | 31.63 | 23.50 | 28.0 | 18.34 | 16.74 | 30.0 | 3.00 | 0.00 | 24.0 | 12.02 | 10.32 | |
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Base | AEGIS | FailBank | |||||||||
| Backbone | Level | Task | SR | CC / | BRS | SR | CC / | BRS | SR | CC / | BRS |
| Safety / Static obstacles: pick the named object and place it in the bowl or plate | |||||||||||
| 1 | T0 Apple | 90.0 | 9.12 / 0.12 | 0.368 | 82.0 | 8.20 / 0.00 | 0.407 | 90.3 | 9.21 / 0.21 | 0.261 | |
| T1 Lemon | 83.0 | 44.07 / 35.87 | 0.368 | 78.0 | 7.81 / 0.01 | 0.524 | 81.0 | 25.98 / 17.88 | 0.446 | ||
| T2 Mango | 77.0 | 74.08 / 66.38 | 0.368 | 68.0 | 21.08 / 14.28 | 0.448 | 93.7 | 23.12 / 13.76 | 0.786 | ||
| T3 Onion | 72.0 | 33.88 / 27.18 | 0.368 | 44.0 | 9.12 / 4.71 | 0.337 | 90.0 | 31.63 / 23.50 | 0.543 | ||
| Task | Base | AEGIS | FailBank |
| L1-T0 Apple | 0.368 | 0.407 [0.109, 0.707] | 0.261 [0.004, 0.535] |
| L1-T1 Lemon | 0.368 | 0.524 [0.299, 0.723] | 0.446 [0.250, 0.619] |
| L1-T2 Mango | 0.368 | 0.448 [0.268, 0.613] | 0.786 [0.657, 0.888] |
| L1-T3 Onion | 0.368 | 0.337 [0.199, 0.459] | 0.543 [0.427, 0.649] |
| L1-T4 Tomato | 0.368 | 0.735 [0.493, 0.926] | 0.571 [0.341, 0.754] |
| L2-T0 Apple | 0.368 | 0.307 [0.243, 0.363] | 0.437 [0.374, 0.492] |
| Task | Base | AEGIS [95% CI] |
|---|---|---|
| L1-T2 Mango | 0.368 | 0.245 [0.009, 0.607] |
| L1-T3 Onion | 0.368 | 0.615 [0.564, 0.663] |
| L2-T0 Apple | 0.368 | 0.402 [0.350, 0.446] |
| L2-T1 Lemon | 0.368 | 0.523 [0.438, 0.609] |
| L2-T2 Mango | 0.368 | 0.515 [0.209, 0.766] |
| L2-T4 Tomato | 0.368 | 0.387 [0.271, 0.502] |
| Model | Arena L1 SR | Action interface | Scope decision |
|---|---|---|---|
| -FFT | 0.64 | Flow matching | Main backbone with complete matched evaluation |
| / -FFT | 0.74 / 0.76 | Flow matching | Second backbone with complete matched evaluation |
| -FAST / -FFT | 0.40 / 0.60 | FAST tokens | Inapplicable because the chunk-level continuous action loss is absent |
| OpenVLA | 0.60 | Discrete autoregressive | Inapplicable to the current loss and guard |
| OpenVLA-OFT | 0.20 | Discrete autoregressive | Inapplicable to the current loss and guard |
| UniVLA | 0.42 | Autoregressive latent action | Inapplicable to the current loss and guard |
| Evidence | Suite / level | Tasks | Protocol |
|---|---|---|---|
| Main static | static obstacles / L1 | T0–T4 | base, AEGIS, FailBank , 50 states |
| State holdout | static obstacles / L1 | T2 | train 32 states, test 18 unseen states |
| Zero-shot static | static obstacles / L2 | T0–T4 | base, AEGIS, 3 FailBank orders |
| Static rounds | static obstacles / L1 | T3 | base and R1–R5, 50 states |
| Dynamic stress test | dynamic obstacles / L1 | T3, T4 | base and learned policies, 18 states |
| Dynamic rounds | dynamic obstacles / L2 | T0 | base and R1–R3, 50 states |
| Radius | Training exposure | Evaluation | Result |
|---|---|---|---|
| Unseen states | L1-T2 training states | held-out L1-T2 states | SR |
| Unseen tasks | L1-T2 only | other Level 1 tasks on | mean SR and paired test in Table 25 |
| Unseen level | Level 1 only | Level 2 T0–T4 | apple/tomato gains, mango/onion preserved, lemon floor |
| Category | Suite | Cost semantics | Count | Measured evidence | Evaluation status |
|---|---|---|---|---|---|
| Safety | Static obstacles | Fall and contact with protected external bodies | 10/15 | Complete matched three-arm results | Main method comparison |
| Safety | Dynamic obstacles | Contact with moving bodies and fall | 15/15 | Base L1-T0 SR 95%, dynamic transfer fails | Negative stress test |
| Safety | Cautious grasp | Part-level gripper distance to target itself | 15/15 | Base SR 0/100. Oracle hazard set empty | Floor and unrepresented constraint |
| Safety | State preservation | Containment predicates | 10/15 | Base SR 66/100. Zero action rewrites in 196 AEGIS chunks | Incompatible avoidance semantics |
| Safety | Hazard avoidance | Surface-distance dwell cost near stove/candle and fall | 15/15 | Initial objects in cost region 42–50/50. VLM misidentifies both preliminary validity cells | Teacher pilot and AEGIS validity gate fail |
| Distractor | Static distractors | No cost predicate | 0/15 | BDDL audit | No benchmark cost axis |
| Level | Task | base | base | ||||
|---|---|---|---|---|---|---|---|
| SR | CC | Logged policy cost | SR | CC | Logged policy cost | ||
| L1 | T0 | 9.0 | 16.75 | 334.9 | 4.0 | 15.99 | 319.7 |
| T1 | 0.0 | 21.55 | 430.6 | 2.0 | 16.68 | 333.2 | |
| T2 | 62.0 | 8.09 | 161.8 | 2.0 | 11.12 | 222.3 | |
| T3 | 34.0 | 15.79 | 315.6 | 12.0 | 20.36 | 406.8 | |
| T4 | 0.0 | 20.67 | 413.5 | 0.0 | 22.40 | 448.0 | |
| Collection | Train records | Guard records | Guard decision | Apple SR |
|---|---|---|---|---|
| Observe-only | 3,657 | 300 | pass | 34.0 |
| Shield in loop | 5,099 | 112 | reject | unavailable |
| Update target | Seed 1 | Seed 2 | Order 3 | Mean |
|---|---|---|---|---|
| Base | 8.0 | 8.0 | 8.0 | 8.0 |
| SFTPOS | 4.0 | 10.7 | 10.7 | 8.4 |
| SFT0 | 18.0 | 14.7 | 25.3 | 19.3 |
| SHAM | 26.0 | 16.0 | 18.0 | 20.0 |
| FailBank | 36.0 | 31.3 | 35.3 | 34.2 |
| Task | Base | FailBank | SFT0 | SHAM | SFTPOS |
|---|---|---|---|---|---|
| L1-T1 Lemon | 83.0 / 35.87 | 83.0 / 12.80 | 84.0 / 5.66 | 94.0 / 7.75 | 81.0 / 29.01 |
| L1-T4 Tomato | 85.0 / 37.01 | 90.0 / 31.24 | 86.0 / 33.51 | 82.0 / 40.61 | 81.0 / 51.27 |
| Task | Method | SR | CC | |
|---|---|---|---|---|
| Apple | Base | 8.0 | 67.51 | 65.89 |
| AEGIS | 7.3 | 90.73 | 89.26 | |
| Teacher in loop | 0.0 | 67.02 | 67.01 | |
| Mango | Base | 89.3 | 51.41 | 33.93 |
| AEGIS | 88.0 | 28.75 | 11.15 | |
| Teacher in loop | 50.7 | 107.63 | 97.50 |
| Method | Updated VLA | Shield | Object geometry | QP/step | s/step |
|---|---|---|---|---|---|
| Base policy | no | no | no | no | 0.362 |
| AEGIS | no | yes | yes | yes | 0.449 |
| FailBank | LoRA | no | no | no | 0.346 |
| Component | Setting |
|---|---|
| PaliGemma 2B LoRA | rank 16, |
| 300M action-expert LoRA | rank 32, |
| Trainable parameters | LoRA factors only |
| Optimizer | AdamW, gradient clipping 1.0 |
| Learning-rate schedule | cosine, 20-step warmup, peak |
| 400-step decay, final |
| Bank | Collection suite | Train records | CBF-triggered / quiet | Validation fold | Used by |
|---|---|---|---|---|---|
| Main static two-round | static L1-T2 | 6,535 | 2,863 / 3,672 | 600 | main tables, L2, R2 ablations |
| Pre-filter static pool | static L1-T2 | 12,343 | 4,485 / 7,858 | 600 | audit only |
| State holdout | static L1-T2 | 4,006 | 1,799 / 2,207 | 600 | unseen-state evaluation |
| Static R1–R5 | static L1-T2 | 3,657 / 6,535 / 15,707 / 19,051 / 22,732 | – | – | round curve |
| Dynamic diagnostic R1 | dynamic L2-T0 | 3,192 | 642 / 2,550 | 139 | dynamic L1 stress test |
| Dynamic diagnostic R2 | dynamic L2-T0 | 9,898 | 1,818 / 8,080 | 266 | dynamic round analysis |
| Bank | Records | Flow-loss ratio | First-action drift | Triggered loss | Decision |
|---|---|---|---|---|---|
| Observe-only round 1 | 3,657 | 0.972 | – | – | accept |
| Shield-in-loop round 1 | 5,099 | 1.137 | 0.0129 | – | reject |
| Round 2 only | 2,878 | 1.111 | 0.0053 | – | reject |
| Accumulated, size matched | 2,878 | 1.008 | – | – | accept |
| Accumulated, full | 6,535 | 1.0065 | 0.00949 | – | accept |
| 0.02 | 0.05 | 0.10 | ||
|---|---|---|---|---|
| 1.05 | 87 | 89 | 89 | 89 |
| 1.10 | 96 | 100 | 100 | 100 |
| 1.15 | 100 | 108 | 108 | 108 |
| 1.20 | 100 | 109 | 109 | 109 |
| 1.30 | 101 | 110 | 110 | 110 |
| L2 Apple | L1 Tomato | |||||
|---|---|---|---|---|---|---|
| Steps | Flow-loss ratio | First-action drift | SR | SR | ||
| Base | – | – | 8.0 | 65.89 | 85.0 | 37.01 |
| 400 | 1.0042 | 0.0083 | 32.0 | 61.71 | 88.0 | 36.65 |
| 800 | 1.0073 | 0.0096 | 36.0 | 65.73 | 90.0 | 31.24 |
| 1,600 | 1.0165 | 0.0108 | 46.7 | 56.92 | 85.0 | 35.71 |
| 2,400 | 1.0341 | 0.0121 | 54.0 | 43.01 | 90.0 | 29.38 |
| Steps | Flow-loss ratio | Triggered loss | Decision | |
|---|---|---|---|---|
| 800 | 0.0 | 1.477 | 0.185 | reject |
| 500 | 0.0 | 1.513 | 0.171 | reject |
| 400 | 0.0 | 1.580 | 0.170 | reject |
| 800 | 0.3 | 1.110 | 0.199 | reject |
| 800 | 0.5 | 1.083 | 0.178 | accept |
| 800, | 0.0 | 1.017 | – | accept |
| Task | Base | |||
|---|---|---|---|---|
| L1-T0 | 54.0 | 58.0 | 34.0 | 30.0 |
| L1-T1 | 66.0 | 84.0 | 60.0 | 62.0 |
| L1-T2 | 92.0 | 96.0 | 94.0 | 90.0 |
| L1-T3 | 28.0 | 26.0 | 24.0 | 14.0 |
| L1-T4 | 56.0 | 88.0 | 78.0 | 78.0 |
| Pooled SR and policy-induced CC | 59.2 / 3.78 | 70.4 / 2.09 | 58.0 / 2.23 | 54.8 / 2.94 |
| Perturbation | Bank size | Corrected targets | Affected share, % | weight mass, % |
|---|---|---|---|---|
| Default | 6,535 | 709 | 0.0 | 0.0 |
| Risk cutoff | 7,683 | 457 | 21.9 | |
| Risk cutoff | 4,892 | 1,293 | 35.6 | |
| Stage boundaries 24 / 12 | 6,626 | 744 | 1.4 | |
| Stage boundaries 36 / 18 | 6,459 | 682 | 1.2 | |
| bins | 6,535 | 709 | 19.2 |
| Task | Base | Observe-only R1 | In-loop R1 | Accumulated matched | R2 only | Full accumulated |
|---|---|---|---|---|---|---|
| L2-T0 Apple | 8.0 / 65.89 | 34.0 / 59.97 | 18.0 / 57.68 | 41.3 / 48.07 | 38.7 / 56.59 | 36.0 / 65.73 |
| L1-T1 Lemon | 83.0 / 35.87 | 85.0 / 24.00 | 93.0 / 3.24 | 89.0 / 10.61 | 87.0 / 8.51 | 83.0 / 12.80 |
| L1-T4 Tomato | 85.0 / 37.01 | 89.0 / 30.57 | 87.0 / 30.55 | 87.0 / 38.53 | 89.0 / 31.23 | 90.0 / 31.24 |
| Task | Base | AEGIS | FailBank | FailBank + AEGIS |
|---|---|---|---|---|
| L2-T2 Mango | 89.3 / 51.41 | 88.0 / 28.75 | 94.0 / 36.78 | 86.0 / 19.78 |
| L2-T3 Onion | 81.3 / 16.33 | 1.3 / 0.27 | 84.7 / 16.93 | 2.7 / 0.53 |
| Analysis | Comparison | Favorable / adverse / tie | -value |
|---|---|---|---|
| Breadth | Static L1-T0 SR | 7 / 7 / 36 | 1.0000 |
| Static L1-T1 SR | 12 / 10 / 28 | 0.8320 | |
| Static L1-T3 SR | 17 / 2 / 31 | 0.0007 | |
| Static L1-T4 SR | 11 / 7 / 32 | 0.4810 | |
| Four unseen L1, SR | 51 / 45 / 104 | 0.61 | |
| Four unseen tasks, legacy policy-induced CC | 71 / 53 / 76 | 0.1265 |
| Backbone | Set | Sign-test | SR | 95% CI | 95% CI | |
|---|---|---|---|---|---|---|
| L1 unseen four | 0.61 | +6.2 | [2.0, 10.5] | -8.39 | [-15.21, -1.97] | |
| L1 all five | 0.099 | +8.3 | [4.5, 12.3] | -17.24 | [-24.22, -10.61] | |
| L2 all five | +8.8 | [6.2, 11.6] | -16.78 | [-24.16, -9.40] | ||
| All ten | +8.6 | [6.2, 10.9] | -17.01 | [-22.19, -12.00] | ||
| L1 unseen four | 0.014 | +11.8 | [5.0, 18.8] | -1.59 | [-3.95, 0.53] | |
| L1 all five | 0.024 | +9.6 | [3.8, 15.6] | -1.08 | [-2.97, 0.62] |
| Task | Training | SR higher / lower / equal | SR -value | Activation higher / lower / equal | Activation -value |
|---|---|---|---|---|---|
| order | |||||
| Apple | 1 | 28 / 2 / 20 | 32 / 18 / 0 | 0.0649 | |
| 2 | 24 / 3 / 23 | 22 / 26 / 2 | 0.6655 | ||
| 3 | 29 / 3 / 18 | 30 / 20 / 0 | 0.2026 | ||
| Mango | 1 | 8 / 4 / 38 | 0.3877 | 14 / 36 / 0 | 0.0026 |
| 2 | 10 / 3 / 37 | 0.0923 | 14 / 36 / 0 | 0.0026 |
| Task | Arm | SR | CC | SR test relative to base | |
| L2-T1 Lemon | Base | 0.0 | 158.14 | 158.13 | – |
| AEGIS | 0.0 | 141.51 | 141.51 | – | |
| FailBank mean | 0.0 | 110.14 | 110.12 | floor | |
| L2-T4 Tomato | Base | 74.0 | 67.87 | 53.07 | – |
| AEGIS | 28.7 | 26.93 | 21.20 | ||
| FailBank order 1 | 84.0 | 59.05 | 42.39 |
| Task | Cost | FailBank -1 | FailBank -2 | FailBank -3 | AEGIS |
|---|---|---|---|---|---|
| Apple | 25/25/0, 1.000 | 25/25/0, 1.000 | 27/23/0, 0.672 | 19/27/4, 0.302 | |
| CC | 22/28/0, 0.480 | 23/27/0, 0.672 | 26/24/0, 0.888 | 18/27/5, 0.230 | |
| Mango | 28/20/2, 0.312 | 34/16/0, 0.015 | 28/20/2, 0.312 | 42/8/0, | |
| CC | 29/19/2, 0.193 | 35/15/0, 0.0066 | 29/21/0, 0.322 | 42/8/0, | |
| Onion | 1/0/49, 1.000 | 1/1/48, 1.000 | 1/0/49, 1.000 | 1/0/49, 1.000 | |
| CC | 9/13/28, 0.524 | 11/11/28, 1.000 | 10/11/29, 1.000 | 49/0/1, |