A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies
Organizations: Huawei Heisenberg Research Center · Huawei Noah’s Ark Lab · TU Berlin · UCL Center for AI
Abstract
A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves open. To bring those futures into the decision, we derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completion. The derivation reveals a candidate-dependent feasible-future mass: its support records whether safe completion remains possible under the frozen continuation process, while its magnitude measures how much weighted safe-completion mass remains. Since exact evaluation is impractical online, we develop a selective finite-candidate approximation and establish conditions for recovering the best retained viable candidate. Our alarm-triggered, training-free reranker VICS-G lowers mean cumulative safety cost by 1.9%-57.5% across six Safety-CHORES settings while remaining within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length. Our approach offers a promising and practical path toward safer task completion, grounded in an exact policy-relative target yet requiring neither policy retraining nor online rollouts.
Figures & tables
| Setting | Decoder | Success | S (pp) | Cost | Cost red. | Viol. | Steps |
|---|---|---|---|---|---|---|---|
| PickUp / base | Policy sample | .894 (143/160) | – | 1.656 | – | 1.656 | 41.78 |
| RCD | .775 (124/160) | .931 | 43.8% | .906 | 33.66 | ||
| VICS-G | .894 (143/160) | 1.625 | 1.9% | 1.625 | 41.83 | ||
| VICS-S | .900 (144/160) | 1.175 | 29.1% | 1.175 | 40.09 | ||
| PickUp / safe | Policy sample | .925 (148/160) | – | .400 | – | .400 | 45.19 |
| RCD | .850 (136/160) | .256 | 36.0% | .244 | 40.76 |
| Decoder | Success (%) | S (pp) | Safe success (%) | Cost | Cost red. | Viol. | Steps |
|---|---|---|---|---|---|---|---|
| Policy sample | 55.09 | – | 28.63 | 4.33 | – | 3.42 | 116.9 |
| RCD † | 37.79 | 23.11 | 1.98 | 54.3% | 1.53 | 104.7 | |
| Switch after failure † | 53.92 | 27.76 | 4.27 | 1.3% | 3.39 | 118.5 | |
| VICS-G (feedback masked) | 55.09 | 28.49 | 4.03 | 6.8% | 3.18 | 117.0 | |
| VICS-G + execution feedback | 55.67 | 28.92 | 3.64 | 15.9% | 2.85 | 116.9 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Support rule | Magnitude | Selected | Safe success | Regret |
|---|---|---|---|---|
| No rejection | Flat | |||
| No rejection | Exact | |||
| Exact support | Flat | |||
| Exact support | Exact |
| Action group | Motions or events | ||||
|---|---|---|---|---|---|
| Forward | Move forward | ||||
| Turn | Rotate left or right | ||||
| Arm translation | Raise, lower, extend, or retract | ||||
| Backward | Move backward | ||||
| Other | Grasp, release, wrist rotation, or termination |
| Setting | Success (pp) | Cost | Violations | Steps |
|---|---|---|---|---|
| PickUp / base | ||||
| PickUp / safe | ||||
| ObjectNav / base | ||||
| ObjectNav / safe | ||||
| Fetch / base | ||||
| Fetch / safe |
| Method | Task success (%) | Safe success (%) | Cost | Steps |
|---|---|---|---|---|
| Policy sample | 56.98 | 28.49 | 4.052 | 118.89 |
| RCD | 40.12 | 28.49 | 1.436 | 106.11 |
| VICS-G | 58.72 | 27.91 | 3.233 | 117.65 |
| VICS-R | 58.72 | 30.23 | 3.116 | 118.34 |
| Decoder | Success (%) | Safe success (%) | Safety cost | Violations | Steps |
|---|---|---|---|---|---|
| Clean † ( executions per method) | |||||
| Policy sample | 58.14 | 30.81 | 3.628 | 2.773 | 116.63 |
| RCD | 40.70 | 25.58 | 1.797 | 1.407 | 104.34 |
| Switch after failure | 58.72 | 30.81 | 3.058 | 2.541 | 116.50 |
| VICS-G (feedback masked) | 59.88 | 31.98 | 3.151 | 2.535 | 115.45 |
| VICS-G + execution feedback | 59.30 | 31.40 | 3.529 | 2.797 | 115.65 |
| Outcome | Clean † | IID shift | Persistent shift |
|---|---|---|---|
| Success (pp) | |||
| Safe success (pp) | |||
| Safety cost | |||
| Violations | |||
| Steps |
| Step / outcome | Means: feedback / ref. | Difference [ CI] | Required bound | Decision |
|---|---|---|---|---|
| A. Primary tests: reference is the same decoder without feedback | ||||
| 1 / Success | pp | Lower pp | Pass | |
| 2 / Safety cost | Upper | Not established; stop | ||
| B. Descriptive comparisons: reference is policy sampling | ||||
| 3 / Success | pp | Lower pp | Not tested | |
| 4 / Safety cost | Upper | Not tested | ||