CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding
Authors: Mohammad Khoshnazar, Mohammad Dehghani Tezerjani, Deyuan Qu, Zhiyuan Gao, Yanxiang Zhan, Jeroen Schafer, Andrew Melnik, Qing Yang, +1 more
Organizations: University of Bremen, Bremen, Germany. · University of North Texas, Denton, TX, USA. · Toyota Motor North America, InfoTech Labs, Mountain View, CA, USA.
Vision-language-action (VLA) policies assume the embodiment on which they were trained and can fail when a joint fault changes how commanded actions are physically executed. Existing fault-recovery methods often require task-specific retraining, fault labels, explicit diagnosis, or privileged embodiment information. We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning. CAPABLE infers capability, how much of the commanded motion each joint actually realizes and how that motion contributes to end-effector behavior, online from command-response history and kinematics using a temporal encoder shared across joints, Jacobian grounding, cross-joint attention, and self-supervised physical prediction. The resulting representation conditions a residual policy that adds bounded corrections to the VLA arm action without fault labels or faulty-joint identifiers. Across 28 LIBERO tasks, CAPABLE raises success on an actuator excluded from fault training from 24.8% to 59.3%, outperforming a parameter-matched global-history baseline by 17.4 points while preserving healthy performance. Leave-one-actuator-out experiments across six joints show that this transfer is not specific to one actuator, and additional evaluations characterize transfer to unseen fault families and demonstrate recovery on a physical Franka Panda. https://capable-vla.github.io/
Figures & tables
Fig. 1: The illustrated sequence (1) opens the drawer , (2) lifts the plate , and (3) places it inside . A faulty joint (red highlight) prevents the frozen VLA motion from succeeding (red cross); CAPABLE supplies a corrective motion (green check). CAPABLE infers each joint’s remaining capability from its command–response history and the live Jacobian, and adds a bounded residual correction to the frozen VLA action. No fault label or explicit joint-index embedding is supplied, allowing the shared estimator to transfer to an actuator whose failure was not observed during training.
Fig. 2: CAPABLE architecture. Shared encoders map per-joint command–response histories and live kinematics to ztcap , which FiLM-conditions the residual SAC actor. The bounded arm correction is added to the frozen VLA action; the gripper bypasses RL. Critics, replay and rewards are training-only.
Method
Base VLA
Residual
Oracle fault
Oracle task state
Joint-fact.
Base VLA (no correction)
Yes
No
No
No
No
POMDP-RL
No
No
No
No
No
Classical RR
Yes
No
No
No
Yes
DEFT
No
No
Yes
Yes
No
Global-history SAC
Yes
Yes
No
No
No
CAPABLE (ours)
Yes
Yes
No
No
Yes
TABLE I: Comparison assumptions. Learned methods use the same task pool, fault conditions, budget, and held-out states. Classical RR is non-learning; DEFT receives simulator-derived goals and actuator availability.
Method
H
S
U
Δ vs. Base
Base VLA
91.4±0.0
41.6±0.0
24.8±0.0
–
POMDP-RL
19.6±2.4
11.3±2.1
10.1±1.9
−14.7
Classical RR
81.6±0.0
61.3±0.0
32.9±0.0
+8.1
DEFT
89.7±1.4
70.1±2.3
62.8±2.9
+38.0
Global-history SAC
88.1±1.3
73.8±2.1
41.9±3.2
+17.1
CAPABLE
90.8±1.0
86.7±1.8
59.3±3.0
+34.5
TABLE II: Main 28-task success (%). H/S/U : healthy/seen/unseen- j2 ; Δ is versus Base VLA on unseen j2 . Classical RR is non-learning; DEFT uses oracle fault and task state.
Suite
Base VLA
Global-history SAC
CAPABLE
Δ vs. Global
LIBERO-10
8.0±0.0
26.5±4.6
45.2±4.1
+18.7
LIBERO-Goal
40.0±0.0
49.0±3.8
73.3±4.7
+24.3
LIBERO-Object
20.8±0.0
52.8±3.3
65.6±4.4
+12.8
LIBERO-Spatial
41.4±0.0
49.5±1.2
64.0±3.8
+14.5
Task-balanced mean
24.8±0.0
41.9±3.2
59.3±3.0
+17.4
TABLE III: Unseen- j2 success by suite (%).
Method
Persistent lock
Damping
Friction
Partial effectiveness
Range restriction
Late-onset lock
Base VLA
24.8±0.0
45.8±0.0
40.6±0.0
44.2±0.0
32.7±0.0
50.4±0.0
Global-history SAC
41.9±3.2
55.6±1.4
51.8±1.6
64.1±3.3
54.7±2.8
61.5±3.1
CAPABLE
59.3±3.0
67.8±2.9
64.2±3.2
62.3±3.0
48.2±3.5
72.4±2.7
Δ vs. Global
+17.4
+12.2
+12.4
−1.8
−6.5
+10.9
TABLE IV: Zero-shot cross-fault success on unseen j2 (%). Training uses only persistent locks; non-lock columns average three severities.
Held-out joint
Global-history SAC
CAPABLE
Δ vs. Global
j0
55.0±4.4
76.7±3.3
+21.7
j1
39.4±1.3
56.2±2.3
+16.8
j2
46.6±5.2
69.3±3.0
+22.7
j3
39.6±4.6
63.6±4.7
+24.0
j4
44.7±5.1
88.1±1.9
+43.4
j5
29.3±4.2
48.2±3.6
+18.9
TABLE V: Leave-one-actuator-out transfer across six held-out joints on eight LIBERO tasks. Each row uses a separate leave-one-actuator-out training split in which the indicated joint is excluded from fault training and evaluated zero-shot. Values are success rates (%).
Variant
P
H
S
U
Δ vs. Full
Full CAPABLE
0.82
90.8±1.0
86.7±1.8
59.3±3.0
–
Global-history SAC
0.82
88.1±1.3
73.8±2.1
41.9±3.2
−17.4
Equal-strata replay
0.82
92.3±0.9
71.5±2.6
46.4±3.7
−12.9
No Jacobian query
0.81
89.8±1.4
75.9±2.4
47.8±3.8
−11.5
No cross-joint attention
0.71
89.2±1.5
81.2±2.7
46.1±3.6
−13.2
No capability loss
0.80
88.7±1.7
72.0±3.1
43.7±4.6
−15.6
TABLE VI: 28-task ablations. P : parameters (M); H/S/U : healthy/seen/unseen- j2 ; Δ is versus full CAPABLE on unseen j2 .
Task
Locked joint
Global-history SAC
CAPABLE
Open drawer
j0 (seen)
5/10 (50%)
9/10 (90%)
Put the object in the drawer
j2 (unseen)
4/10 (40%)
7/10 (70%)
Object rearrangement
j6 (seen)
4/10 (40%)
8/10 (80%)
Overall
–
13/30 (43.3%)
24/30 (80.0%)
TABLE VII: Real-robot success under persistent locks, 10 trials per condition. Locks are software-enforced at the joint-reference layer.