Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models
Organizations: Showlab, National University of Singapore
Abstract
Vision-language-action (VLA) models adapted through supervised fine-tuning (SFT) inherit a structural asymmetry: expert demonstrations teach the policy where success behavior lies, but provide no signal about where it ceases to be reliable. We argue that robust VLA adaptation should therefore be viewed not as further demonstration fitting, but as Failure-Boundary Learning -- the problem of Discovering, Localizing, and Shaping the boundary between recoverable deviations and task failure. To instantiate this view, we propose DLS: built on a real-grounded behavioral prior from few real demonstrations and simulated co-training, DLS discovers failure boundaries at scale through on-policy digital twin rollouts. Rather than reducing each rollout to a binary label, semantic progress localization uses privileged simulator states to assign progress-aware signals that capture where the failure boundary is crossed, not merely whether. These signals drive directional boundary shaping in the flow dynamics -- reinforcing success-producing denoising directions and suppressing failure-producing ones, without action likelihoods or auxiliary critics. Across real-robot manipulation tasks, DLS improves robustness over SFT and online RL baselines, especially under randomized initial states and unseen visual conditions.
Figures & tables
| Model | Training Setting | Pick & Place | Sweep | Plugin | Avg. ( ) | ||||
| ID. | Unseen. | ID. | Unseen. | ID. | Unseen. | ID. | Unseen. | ||
| # Full SFT | |||||||||
| Real-only ( N = 50 ) | 20/30 | 18/30 | 16/30 | 14/30 | 22/30 | 13/30 | 64.4 [54.1, 73.6] | 50.0 [39.9, 60.1] | |
| Real-only ( N = 100 ) | 25/30 | 19/30 | 20/30 | 15/30 | 24/30 | 15/30 | 76.7 [66.9, 84.2] | 54.4 [44.2, 64.3] | |
| # Few-shot SFT + RL | |||||||||
| Sim-Real SFT ( N = 50 ) | 20/30 | 15/30 | 13/30 | 12/30 | 17/30 | 9/30 | 55.6 [45.3, 65.4] | 40.0 [30.5, 50.3] | |
| Method | Pick & Place | Sweep | Plugin | |||
| Unseen SR. ( ) | Hor. ( ) | Unseen SR. ( ) | Hor. ( ) | Unseen SR. ( ) | Hor. ( ) | |
| Step-count Penalty | 279.5 | 296.1 | 283.9 | |||
| Distance-based Potential | 287.5 | 301.4 | 292.0 | |||
| SPL (ours) | 280.3 | 300.1 | 287.4 | |||
| Signal | Agree. (%) | Unseen SR | ||
| Majority phase | 52.1 | 0.02 | 41 | — |
| Step-count | 41.2 | 0.18 | 62 | 56.7 |
| Distance potential | 46.4 | 0.31 | 34 | 60.0 |
| SPL (ours) | 78.1 | 0.68 | 19 | 72.2 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Phase | Gate Predicate ( or Success) | |||
| Pick & Place [2pt] :Block :Coaster | 1 | 0.10 | Reach | ||
| 2 | 0.25 | Grasp | |||
| 3 | 0.20 | Transit | G | ||
| 4 | 0.45 | Place | |||
| Sweep [2pt] :Broom :Cube :Dustpan | 1 | 0.08 | Reach | ||
| 2 | 0.20 | Grasp |
| Parameter | Pick & Place | Sweep | Plugin |
| Model Config | |||
| Denoising steps | 5 | 5 | 5 |
| Action chunk / horizon | 8 / 8 | 8 / 12 | 8 / 8 |
| Control frequency | 10 Hz (stride 2) | 10 Hz (stride 2) | 10 Hz (stride 2) |
| Sim-Real SFT | |||
| Base model | lerobot/pi05_base | lerobot/pi05_base | lerobot/pi05_base |
| Training Setting | Pick & Place | Sweep | Plugin | Avg. | ||||
| ID. | Unseen. | ID. | Unseen. | ID. | Unseen. | ID. | Unseen. | |
| Real-only ( N = 50 ) | 48.8–80.8 | 42.3–75.4 | 36.1–69.8 | 30.2–63.9 | 55.6–85.8 | 27.4–60.8 | 54.1–73.6 | 39.9–60.1 |
| Real-only ( N = 100 ) | 66.4–92.7 | 45.5–78.1 | 48.8–80.8 | 33.2–66.8 | 62.7–90.5 | 33.2–66.8 | 66.9–84.2 | 44.2–64.3 |
| Sim-Real SFT ( N = 50 ) | 48.8–80.8 | 33.2–66.8 | 27.4–60.8 | 24.6–57.7 | 39.2–72.6 | 16.7–47.9 | 45.3–65.4 | 30.5–50.3 |
| + GRPO [ 29 ] | 55.6–85.8 | 52.1–83.3 | 42.3–75.4 | 27.4–60.8 | 62.7–90.5 | 42.3–75.4 | 61.0–79.5 | 47.5–67.5 |
| Variant | Agree. | Sim. SR |
| : 12 Dirichlet draws | — | 78.0 5.4 |
| : / | — | 77.5 / 74.8 |
| : order reversed | — | 63.3 |
| Gate thresholds | 66.2 | 75.2 |
| Merged phases ( ) | 44.7 | 71.0 |
| SPL (ours) / Terminal-only | 78.1 / — | 79.4 / 67.7 |