Vision-Language-Action (VLA) models rely strongly on language for describing task information, despite having multimodal inputs. We hypothesize that other modalities in the state space may present opportunities for supplemental task conditioning, which may be particularly relevant in cluttered or otherwise ambiguous scenes. We introduce two tuned models to test this hypothesis: (1) an electrophysiology-conditioned VLA (EC-VLA) that incorporates 8-channel electromyography envelopes as continuous conditioning input concatenated to the proprioceptive vector, and (2) a visually-annotated VLA (VA-VLA) that incorporates visual segmentation annotations to the image inputs. On a cube-selection task evaluated across three participants, EC-VLA matches a language-prompted baseline in uncluttered, in-distribution conditions and substantially outperforms it in cluttered, out-of-distribution scenes. Similarly, VA-VLA shows modest improvements over a language-prompted baseline in in-distribution scenes with substantial improvement in cluttered, out-of-distribution trials. Together, these results provide strong evidence for the potential benefit of task-conditioning beyond language.
Figures & tables
Fig. 1: An overview of vision-language-action (VLA) model architectures and how we propose to modify the inclusion of task information.
Completion
Criteria
Criteria
Stage
Score
(EC-VLA)
(VA-VLA)
3
1.0
Success
Success
2
0.75
Touched target
Lifted target
1
0.5
Moved toward target
Moved target
0
0.0
Failure
Failure
TABLE I: Task completion scoring system.
Fig. 2: Camera views and gestures from the first experiment. Top row: Visual state information from the scene, showing an overhead view of four cubes in an uncluttered setting on the left, a typical wrist view in the center, and an example of a cluttered scene on the right. Bottom row: Wrist gestures used by participants to condition EC-VLA during training and test.
Fig. 3: Behavior detail for the EC-VLA (left) and VA-VLA (right) trials, showing various partial success criteria as portions of trials. These include failure (red), stage 1 partial credit (yellow), stage 2 partial credit (light green), and success (dark green). The black lines show changes in task completion score, which is an aggregate score of these behaviors. EC-VLA values are averages of all three participant models, and VA-VLA values are only taken from the tasks that can add clutter. (C) denotes cluttered trials
Model
Completion
Success
Stage 2
Stage 1
Failure
Evaluated
Score
Rate
Rate
Rate
Rate
Experiment 1
SmolVLA
0.91
0.80
0.94
0.94
0.06
EC-VLA
0.94
0.80
0.97
1.0
0.00
Delta
+0.03
0.00
+0.03
+0.06
-0.06
SmolVLA (C) 1 1 1 (C) notes cluttered trials. For experiment 2, (UC) notes uncluttered results for the subset of all tasks that were tested further with clutter.
0.43
0.30
0.40
0.50
0.50
TABLE II: Evaluation Results between tuned baselines and our augmented models.
Fig. 4: Details of performance for each participant EC-VLA model. A: partial scoring chart for each model compared with the baseline, showing score 0 (Failure, red), score 0.5 (Stage 1, orange), score 0.75 (Stage 2, light green), and score 1.0 (Success, dark green). B: analysis of how often each model attempted to recover from a failed grasp, and whether those attempts were successful.
Fig. 5: State information used in experiment 2, which always uses a left camera, a wrist camera, and a right camera. In the baseline model, the images are unaltered and paired with a text description (top). In VA-VLA (center), the text is generic and the left and right cameras have visual annotations for the target object (green) and destination (blue). In the cluttered trials (bottom), additional objects are placed in the scene. In practice this uses all three views and is either annotated or not depending on the model.
Task
Lang
Mask
Lang+Clutter
Mask+Clutter
Clutterable Tasks
Cabinet-to-Counter
0.22
0.40
0.14
0.24
Counter-to-Blender
0.20
0.18
0.02
0.02
Counter-to-Cabinet
0.38
0.38
0.06
0.26
Counter-to-Drawer
0.30
0.44
0.14
0.30
Counter-to-Sink
0.30
0.28
0.18
0.16
TABLE III: Success rates for all 16 tasks in experiment 2.
Task
Lang
Mask
Lang+Clutter
Mask+Clutter
Clutterable Tasks
Cabinet-to-Counter
0.470
0.650
0.330
0.510
Counter-to-Blender
0.605
0.595
0.255
0.330
Counter-to-Cabinet
0.595
0.645
0.255
0.525
Counter-to-Drawer
0.625
0.780
0.405
0.625
Counter-to-Sink
0.530
0.625
0.375
0.510
TABLE IV: Task completion scores for all 16 tasks in experiment 2.