CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding
Authors: Mohammad Khoshnazar, Mohammad Dehghani Tezerjani, Deyuan Qu, Zhiyuan Gao, Yanxiang Zhan, Jeroen Schafer, Andrew Melnik, Qing Yang, +1 more
Organizations: University of Bremen, Bremen, Germany. · University of North Texas, Denton, TX, USA. · Toyota Motor North America, InfoTech Labs, Mountain View, CA, USA.
Vision-language-action (VLA) policies assume the embodiment on which they were trained and can fail when a joint fault changes how commanded actions are physically executed. Existing fault-recovery methods often require task-specific retraining, fault labels, explicit diagnosis, or privileged embodiment information. We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning. CAPABLE infers capability, how much of the commanded motion each joint actually realizes and how that motion contributes to end-effector behavior, online from command-response history and kinematics using a temporal encoder shared across joints, Jacobian grounding, cross-joint attention, and self-supervised physical prediction. The resulting representation conditions a residual policy that adds bounded corrections to the VLA arm action without fault labels or faulty-joint identifiers. Across 28 LIBERO tasks, CAPABLE raises success on an actuator excluded from fault training from 24.8% to 59.3%, outperforming a parameter-matched global-history baseline by 17.4 points while preserving healthy performance. Leave-one-actuator-out experiments across six joints show that this transfer is not specific to one actuator, and additional evaluations characterize transfer to unseen fault families and demonstrate recovery on a physical Franka Panda. https://capable-vla.github.io/
Figures & tables
Fig. 1: The illustrated sequence (1) opens the drawer , (2) lifts the plate , and (3) places it inside . A faulty joint (red highlight) prevents the frozen VLA motion from succeeding (red cross); CAPABLE supplies a corrective motion (green check). CAPABLE infers each joint’s remaining capability from its command–response history and the live Jacobian, and adds a bounded residual correction to the frozen VLA action. No fault label or explicit joint-index embedding is supplied, allowing the shared estimator to transfer to an actuator whose failure was not observed during training.
Fig. 2: CAPABLE architecture. Shared encoders map per-joint command–response histories and live kinematics to ztcap , which FiLM-conditions the residual SAC actor. The bounded arm correction is added to the frozen VLA action; the gripper bypasses RL. Critics, replay and rewards are training-only.
Method
Base VLA
Residual
Oracle fault
Oracle task state
Joint-fact.
Base VLA (no correction)
Yes
No
No
No
No
POMDP-RL
No
No
No
No
No
Classical RR
Yes
No
No
No
Yes
DEFT
No
No
Yes
Yes
No
Global-history SAC
Yes
Yes
No
No
No
CAPABLE (ours)
Yes
Yes
No
No
Yes
TABLE I: Comparison assumptions. Learned methods use the same task pool, fault conditions, budget, and held-out states. Classical RR is non-learning; DEFT receives simulator-derived goals and actuator availability.
Method
H
S
U
Δ vs. Base
Base VLA
91.4±0.0
41.6±0.0
24.8±0.0
–
POMDP-RL
19.6±2.4
11.3±2.1
10.1±1.9
−14.7
Classical RR
81.6±0.0
61.3±0.0
32.9±0.0
+8.1
DEFT
89.7±1.4
70.1±2.3
62.8±2.9
+38.0
Global-history SAC
88.1±1.3
73.8±2.1
41.9±3.2
+17.1
CAPABLE
90.8±1.0
86.7±1.8
59.3±3.0
+34.5
TABLE II: Main 28-task success (%). H/S/U : healthy/seen/unseen- j2 ; Δ is versus Base VLA on unseen j2 . Classical RR is non-learning; DEFT uses oracle fault and task state.
Suite
Base VLA
Global-history SAC
CAPABLE
Δ vs. Global
LIBERO-10
8.0±0.0
26.5±4.6
45.2±4.1
+18.7
LIBERO-Goal
40.0±0.0
49.0±3.8
73.3±4.7
+24.3
LIBERO-Object
20.8±0.0
52.8±3.3
65.6±4.4
+12.8
LIBERO-Spatial
41.4±0.0
49.5±1.2
64.0±3.8
+14.5
Task-balanced mean
24.8±0.0
41.9±3.2
59.3±3.0
+17.4
TABLE III: Unseen- j2 success by suite (%).
Method
Persistent lock
Damping
Friction
Partial effectiveness
Range restriction
Late-onset lock
Base VLA
24.8±0.0
45.8±0.0
40.6±0.0
44.2±0.0
32.7±0.0
50.4±0.0
Global-history SAC
41.9±3.2
55.6±1.4
51.8±1.6
64.1±3.3
54.7±2.8
61.5±3.1
CAPABLE
59.3±3.0
67.8±2.9
64.2±3.2
62.3±3.0
48.2±3.5
72.4±2.7
Δ vs. Global
+17.4
+12.2
+12.4
−1.8
−6.5
+10.9
TABLE IV: Zero-shot cross-fault success on unseen j2 (%). Training uses only persistent locks; non-lock columns average three severities.
Held-out joint
Global-history SAC
CAPABLE
Δ vs. Global
j0
55.0±4.4
76.7±3.3
+21.7
j1
39.4±1.3
56.2±2.3
+16.8
j2
46.6±5.2
69.3±3.0
+22.7
j3
39.6±4.6
63.6±4.7
+24.0
j4
44.7±5.1
88.1±1.9
+43.4
j5
29.3±4.2
48.2±3.6
+18.9
TABLE V: Leave-one-actuator-out transfer across six held-out joints on eight LIBERO tasks. Each row uses a separate leave-one-actuator-out training split in which the indicated joint is excluded from fault training and evaluated zero-shot. Values are success rates (%).
Variant
P
H
S
U
Δ vs. Full
Full CAPABLE
0.82
90.8±1.0
86.7±1.8
59.3±3.0
–
Global-history SAC
0.82
88.1±1.3
73.8±2.1
41.9±3.2
−17.4
Equal-strata replay
0.82
92.3±0.9
71.5±2.6
46.4±3.7
−12.9
No Jacobian query
0.81
89.8±1.4
75.9±2.4
47.8±3.8
−11.5
No cross-joint attention
0.71
89.2±1.5
81.2±2.7
46.1±3.6
−13.2
No capability loss
0.80
88.7±1.7
72.0±3.1
43.7±4.6
−15.6
TABLE VI: 28-task ablations. P : parameters (M); H/S/U : healthy/seen/unseen- j2 ; Δ is versus full CAPABLE on unseen j2 .
Task
Locked joint
Global-history SAC
CAPABLE
Open drawer
j0 (seen)
5/10 (50%)
9/10 (90%)
Put the object in the drawer
j2 (unseen)
4/10 (40%)
7/10 (70%)
Object rearrangement
j6 (seen)
4/10 (40%)
8/10 (80%)
Overall
–
13/30 (43.3%)
24/30 (80.0%)
TABLE VII: Real-robot success under persistent locks, 10 trials per condition. Locks are software-enforced at the joint-reference layer.
Deploying Vision-Language-Action (VLA) models in real robotic systems requires robustness not only to semantic and perceptual variations, but also to embodiment-side faults that change how actions are physically realized. Real robots can experience joint-level changes caused by actuator degradation, hardware faults, safety limits, collision damage, or wear-induced friction. These faults are critical because they alter the action-to-motion interface of a policy, disrupting the learned closed-loop relationship between commanded actions, realized motion, and subsequent observations. In this work, we study realistic joint-level physical faults and show that VLA models are vulnerable when predicted actions are executed through a perturbed robot body. Our analysis reveals joint-dependent effects, with heterogeneous degradation in task success across affected joints. We also show that performance drops cannot be attributed solely to physical infeasibility, since feasible faults such as increased joint friction can still substantially reduce success rates and induce closed-loop execution mismatch. Motivated by these findings, we propose Joint-level Physical-fault Aware Residual Calibrator (J-PARC), a lightweight residual calibration framework built on top of a frozen VLA policy. J-PARC infers a latent joint-fault regime from recent joint dynamics and conditions a shared residual calibrator on this regime, enabling adaptive action correction across faulty joints. Experiments show that J-PARC improves robustness under joint-level faults while preserving fault-free environment performance.
Minsoo Jo, Taeju Kwon, Junha Chun +2
Graduate School of Data Science, Seoul National University
Vision-language-action (VLA) policies provide strong priors for language-conditioned manipulation, but remain brittle in off-nominal states requiring targeted recovery. We propose ReCoVLA -- a failure-conditioned residual recovery framework that keeps a pretrained VLA policy frozen, uses an external vision-language model (VLM) to infer the failure mode and recovery stage, and compiles a structured reward from task-relevant components. Rather than using the VLM to generate actions or rewards directly, ReCoVLA uses it as a semantic reward selector: it predicts a recovery descriptor and reward mask for in-simulation residual-policy training, followed by zero-shot sim-to-real deployment of the trained recovery policies. This decouples high-level failure understanding from low-level corrective control to support different VLAs. Experiments across short-horizon, long-horizon, and contact-rich manipulation tasks show that ReCoVLA outperforms the tested baselines on average. In simulation, our reward compiler improves average success from 36.7% for the fine-tuned π0.5 baseline to 66.7%. In physical zero-shot sim-to-real experiments, ReCoVLA achieves the best average performance, with 61.7% success.
Haodi Hu, Chung-Ta Huang, Jing Liu +4
University of Southern California, USA · Mitsubishi Electric Research Laboratories (MERL), USA · Harvard University, USA
Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervision for execution drift, while failed rollouts are often discarded. We introduce RePO-VLA, a recovery-driven policy optimization framework that assigns distinct roles to success, recovery, and failure trajectories. RePO-VLA first applies Recovery-Aware Initialization (RAI), slicing recovery segments and resetting history so corrective actions depend on the current adverse state rather than the preceding failure. It then learns a Progress-Aware Semantic Value Function (PAS-VF), aligning spatiotemporal trajectory features with instructions and successful references. The resulting labels salvage useful failure prefixes via reliability decay, while low-value labels mark drift and terminal breakdowns, teaching differences among nominal, failed, and corrective actions. The data engine turns adverse states into planner-generated or human-collected corrective rollouts, teaching recovery to the success manifold. Value-Conditioned Refinement (VCR) trains the policy to prefer high-progress actions. At deployment, a fixed high value (v=1.0) biases actions toward the learned success manifold without online failure detectors or heuristic retries. We introduce FRBench, with standardized error injection and recovery-focused evaluation. Across simulated and real-world bimanual tasks, RePO-VLA improves robustness, raising adversarial success from 20% to 75% on average and up to 80% in scaled real-world trials.
Weijia Liufu, Xiaoyu Guo, Ruiyi Chen +16
1Sun Yat-sen University · 2South China University of Technology · 3Peng Cheng Laboratory +3