Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires. We call this failure instruction-action binding. Instructions cue familiar trajectory families, and visual feedback adjusts their execution. Behavioral analyses of fine-tuned π0.5 and GR00T-N1.7 policies reveal that failed rollouts often retain the source behavior or switch to another demonstrated task. These switches show that language is not simply ignored. Readouts and interventions connect these choices to task-conditioned internal states. Our analysis of the imitation objective shows how narrow conditional action support can leave grounded and instruction-keyed solutions indistinguishable on the demonstrations. This motivates Equivariant Counterfactual Training (ECT), which acts at two levels. ECT data supply valid demonstrations in which the same instruction requires different actions in distinguishable scenes, while the ECT loss trains each demonstration with its counterpart in the same update. In a controlled LIBERO-PRO comparison, full ECT raises π0.5's mean position-swap success from 36% to 59%. On CALVIN, where counterparts already occur in the original data, the ECT loss improves five-task completion without new demonstrations. On a real UR5e under a fixed demonstration budget, full ECT raises unseen-position success from 8% to 88%.
Figures & tables
Figure 1: Perturbed instructions retrieve memorized trajectories. One LIBERO-Goal task throughout; dashed blue: intended trajectory, purple : the source task’s memorized trajectory, red : another memorized task’s. (a) Nuisance perturbations keep the memorized trajectory valid; counterfactual ones need a new one (Sem.: rephrasing, Obj.: new appearance, Swap: swapped positions, Task: new target). (b) Under a new verb, π0.5 retrieves the source task. (c) Under a new object, it retrieves another memorized task: the instruction selects which trajectory to replay. Of 381 labeled rollouts, 52% replay the source task and 15% another task ( Table 10 ).
Cell
Δk
Δa
Type
Binding pattern
Semantic (Sem.)
0
0
Nuisance
familiar action remains valid
Object (Obj.)
0
0
Nuisance
familiar action can remain valid
Swap
0
1
Counterfactual
same key, different required action
Task
0 or 1
1
Counterfactual
source retention or other-task retrieval
Table 1: Nuisance and counterfactual cells under the binding model. Δk describes a change in the selected instruction key and Δa a change in the required action ( Equation 4 ). The Semantic row represents key-preserving paraphrases. A Task change can retain or switch the key.
Figure 2: Counterfactual failure is often coherent rather than unstructured. Left. Official-checkpoint success on LIBERO-PRO separates nuisance and counterfactual cells. Right. Each pie shows the behavioral labels for one perturbation type (see Table 8 for label definitions). Across all 381 targeted π0.5 probes, 69% are retrieval-like and 4% are unstructured collapse.
Figure 3: Intuition for ECT data and ECT loss. Contours show the original and counterfactual losses. Shading marks a region where both are low. (a) Standard training uses only original demonstrations. (b) ECT data supply alternative supervision. The depicted independent-sampling sequence processes counterparts in separate updates. Independent batches can also contain both branches without matching counterparts. (c) The ECT loss evaluates matched counterparts at the same parameters and combines their gradients in one update. The paths show one illustrative geometry.
Figure 4: ECT scene constructions across the four LIBERO suites. Each group shows an original scene and three mirrored versions under the same instruction. The training construction also uses shifts on Object and selected Long tasks. Table 20 gives the transform families.
Figure 5: Useful counterfactuals must reveal the required action. On Object, front-back mirroring changes the required action but leaves dominant landmarks in similar image regions. An added shift separates the branches visually. Table 19 reports performance by transformation.
Spatial
Object
Goal
Long
Model
Swap
Task
Swap
Task
Swap
Task
Swap
Task
π0.5 Frozen LM (vision encoder and action expert trained)
Standard
52.7
22.3
38.3
23.0
41.3
20.6
12.7
19.5
ECT data + ECT loss
74.4
(+21.7)
22.1
( − 0.2)
70.8
(+32.5)
31.3
(+8.3)
55.5
(+14.2)
35.7
(+15.1)
34.7
(+22.0)
22.5
(+3.0)
π0.5 Full FT (all parameters trained)
Official π0.5
46.6
53.0
18.2
11.0
34.2
21.2
9.4
17.0
Table 2: ECT counterfactual performance across architectures. Success (%). Frozen-LM controls match steps, batch size, and trainable parameters. Full-FT and GR00T use official checkpoints, not matched retrained controls. GR00T ECT models use the ECT data with the standard loss, and four-suite training changes scope ( Section G.2 ). Parentheses give percentage-point gains over Standard (Frozen LM) or the official checkpoint (other blocks). ECT rows are shaded. Frozen-LM rows average three seeds, and other rows are single models.
Spatial
Object
Goal
Long
Standard
52.7
38.3
41.3
12.7
ECT data
72.1
71.1
51.8
29.5
ECT data + ECT loss
74.4
70.8
55.5
34.7
Δdata (ECT data − Standard)
+19.4
+32.8
+10.5
+16.8
Δloss (ECT data + ECT loss − ECT data)
+2.3
−0.3
+3.7
+5.2
Table 3: ECT data and ECT loss make distinct contributions. Swap success (%). All rows use Frozen-LM training with matched steps and batch size (64 examples per update). Each row averages three training seeds ( Section G.3 ). Pairing draws its pairs from the same pool as ECT data. Full five-cell results are in Table 23 . ECT rows are shaded, with the best result per column in bold.
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Instruction-action binding framework
D
Demonstration dataset.
(si,ℓi,ai)
Observation, instruction, and expert trajectory from task i .
T
Task identity.
H(T∣x)
Conditional entropy: remaining uncertainty about task identity T after observing cue x , such as instruction ℓ or observation s .
a⋆(s,ℓ)
Task-level action requirement, abstracting from incidental execution variation.
Appendix
Table 5: Frequently used notation in the main paper and appendix.
Experiment
Evaluation unit
Measurement
LIBERO-PRO diagnosis
Perturbation rollout
Terminal-state simulator predicate
Human rollout labels
Rollout video
Behavioral category
Prefix-KV retrieval
same_to_other rollout
Nearest-neighbor executed-task recovery
Persistent intervention
Rollout
End-effector proximity and simulator predicate
Appendix
Table 6: Evaluation components. All LIBERO success-rate tables use terminal-state success. Human labels and mechanistic diagnostics are separate analyses.
π0.5
GR00T-N1.7
Suite
ID
Sem.
Obj.
Swap
Task
ID
Sem.
Obj.
Swap
Task
Spatial
98.4
98.0
98.0
46.6
53.0
93.6
88.8
93.0
1.2
51.4
Object
98.6
99.0
94.4
18.2
11.0
95.6
97.6
86.0
0.0
9.0
Goal
96.4
94.6
88.8
34.2
21.2
94.2
93.8
76.0
2.2
10.0
Long
91.8
90.8
66.8
9.4
17.0
89.0
88.4
54.6
0.2
10.2
Appendix
Table 7: LIBERO-PRO success of the official checkpoints. π0.5 uses one four-suite checkpoint and GR00T-N1.7 uses per-suite checkpoints. We evaluate N=500 episodes per cell. Fig. 2 (left) shows the four-suite means. Sem. and Task deliver the perturbed instruction to the policy ( Appendix B ).
Figure 6: Goal rollouts can follow familiar training trajectories. Medoid rollouts for the Goal perturbation rules are overlaid on the 80th-percentile envelopes of the corresponding tasks’ 50 training demonstrations. The medoids are selected from 117 rollouts under Rg,1 to Rg,4 . Task 3 groups destinations for readability. The comparisons illustrate whole-task matches.
Fine-grained label
Coarse group
Definition
same_to_origin
Source retrieval
Clearly executes the original source behavior
close_to_origin
Source retrieval
Approximately follows the source behavior
same_to_other
Other-task retrieval
Matches another familiar demonstrated task
same_to_change
Correct or near-correct
Executes the modified instruction correctly
close_to_change
Correct or near-correct
Approximately follows the modified instruction
mix
Hybrid retrieval
Combines source, changed, or neighboring sub-trajectories
Appendix
Table 8: Human rollout annotation taxonomy. Fine labels map to the main paper’s coarse groups.
Granularity
N
Agreement
Cohen’s κ
Macro F1
Fine labels
117
73.5 [65.0, 81.2]
0.611 [0.504, 0.712]
45.9 [35.5, 54.0]
Coarse 5-class
117
83.8 [76.9, 90.6]
0.731 [0.617, 0.833]
58.5 [47.0, 67.6]
Retrieval-like binary
117
90.6 [84.6, 95.7]
0.778 [0.639, 0.894]
88.9 [81.9, 94.7]
Appendix
Table 9: Inter-annotator reliability. 117 of the 381 probes were labeled by both annotators. Brackets give 95% episode-bootstrap confidence intervals.
Suite
Rule
N
Source
Other-task
Hybrid
Correct
Collapse
Goal
Rg,1
33
27
2
0
0
4
Rg,2
33
12
12
0
9
0
Rg,3
21
6
9
0
6
0
Rg,4
30
7
21
0
0
2
Rg,5
3
3
0
0
0
0
total
120
55
44
0
15
6
Appendix
Table 10: Human-labeled outcomes of the official π0.5 by rule and suite. Goal subrules are aggregated into their parent rules. Rs,2 pools two variants of the reworded location.
Trajectory metric
Usable rows
Label-consistent NN
Object+EEF DTW
99
87/99 (87.9%)
EEF+gripper resampled distance
99
83/99 (83.8%)
Appendix
Table 11: Whole-trajectory nearest-neighbor check on Goal. Goal rollouts labeled as retrieving the source or another task, with cached trajectories. A rollout is label-consistent when its nearest expert demonstration belongs to the task the human label names.
Condition
N
Observed behavior
Same instruction, swapped object + 10 cm shift
9
9 / 9 follow the swapped-and-shifted source-bound object
Appendix
Table 12: Object stress test with unchanged language. The instruction is unchanged. The object occupying the source-bound slot is swapped and shifted by 10 cm.
Suite
Rule
Description
Spatial
R1
Replace the target object
R2
Alter the spatial description of the source location (object repositioned to match), with two variants
R3
Replace the placement destination
Object
R1
Swap the target object’s position with another in-scene object
R2
Slightly shift the target object’s position
stress
Swap and shift simultaneously
Appendix
Table 13: Perturbation rules for each LIBERO suite. Rule IDs match the row labels in Fig. 7 . Object rules leave the instruction unchanged and perturb only the scene.
Figure 7: Rollout examples for all four LIBERO suites. Each panel shows the source behavior (Origin), followed by representative rollouts under the perturbation rules of Table 13 , with outcome labels from Table 8 .
Figure 8: A hybrid rollout combining familiar task segments. Under a modified two-object instruction, π0.5 combines a complete segment associated with demonstrated task T1 with a partial segment associated with task T0 .
Source completion
NN source
Suite
Rule
N
Official
ECT
Official
ECT
Goal
Rg,1
33
84.8
66.7
90.9
78.8
Rg,2
33
39.4
9.1
45.5
36.4
Rg,3
21
38.1
19.0
38.1
19.0
Rg,4
30
16.7
0.0
20.0
3.3
total
117
46.2
24.8
50.4
36.8
Appendix
Table 14: Binding probes before and after ECT. Source completion measures how often the original task’s predicate fires under a perturbed instruction. NN source measures how often the nearest expert demonstration belongs to the source task. Values are percentages of rollouts for one ECT seed.
Official π0.5
Official GR00T-N1.7
π0.5 + ECT
Suite
N
Follow
Retr.
Coll.
Follow
Retr.
Coll.
Follow
Retr.
Coll.
Spatial
120
42.5
51.7
5.8
33.3
54.2
12.5
64.2
31.7
4.2
Object
69
40.6
58.0
1.4
15.9
84.1
0.0
75.4
24.6
0.0
Goal
120
12.5
82.5
5.0
0.8
71.7
27.5
26.7
60.8
12.5
Long
72
9.7
87.5
2.8
0.0
97.2
2.8
27.8
72.2
0.0
All
381
26.5
69.3
4.2
13.6
73.2
13.1
47.5
47.2
5.2
Appendix
Table 15: Matched human labels for three policies on the 381 probes. Follow denotes correct or near-correct execution of the changed task. Retr. denotes source-task, other-task, or hybrid retrieval. Coll. denotes behavior without a coherent trajectory. Values are percentages of rollouts. ECT is the Frozen-LM π0.5 model trained with ECT data + ECT loss (seed 42).
Figure 9: Executed tasks are recoverable from scene-matched prefix-KV states. (a) Rank of the executed task among the ten candidates, for scene-matched prefix-KV states, a text-only TF-IDF baseline over the instructions, and a mixed-scene prefix-KV control. (b) The perturbed input is closer to the executed task than to the source task in scene-matched prefix-KV space.
Figure 10: Prefix-KV retrieval is robust to representation aggregation. Same 45 rollouts as Fig. 9 . (a) Rank-1 recovery under alternative representations. Error bars are 95% Wilson intervals, and the dashed line marks chance. (b) Single-layer prefix K+V retrieval by VLM layer.
Patch
Target
First L2 impr.
Pos. frac.
Chunk L2 impr.
Pos. frac.
Source final
Source
0.0024[0.0011,0.0036]
0.69
0.0182[0.0107,0.0262]
0.87
Executed final
Executed
0.0009[−0.0004,0.0022]
0.62
0.0021[0.0006,0.0035]
0.62
Source all
Source
−0.0208[−0.0418,−0.0012]
0.42
0.5467[0.3953,0.6949]
0.84
Executed all
Executed
−0.0094[−0.0237,0.0045]
0.42
0.1048[0.0356,0.1786]
0.62
Appendix
Table 16: Prefix-KV patching intervention. Improvement ΔL2 ( Equation 17 ) with bootstrap intervals over the 45 rollouts. Positive values mean the patched decode is closer to the target expert than the natural perturbed decode. Pos. frac. is the fraction of rollouts with positive improvement.
Analysis
Own target
Off-target
Specificity
Cosine alignment
0.443[0.400,0.487]
0.182[0.148,0.217]
0.262[0.234,0.288]
Normalized projection
0.0136[0.0102,0.0174]
0.0064[0.0044,0.0085]
0.0073[0.0054,0.0092]
Raw projection
0.0215[0.0160,0.0275]
0.0089[0.0061,0.0120]
0.0125[0.0096,0.0158]
Appendix
Table 17: Target specificity of prefix-KV patching. Own target measures alignment of the patch-induced displacement with the patched candidate’s expert trajectory. Off-target averages alignment with the other candidates. Specificity is their difference. Brackets report rollout-bootstrap intervals.
Model
Injected closer
Control closer
Pairs, larger shift
Pairs 4/4
Pairs 0/4
Control completes source
Injected completes target
π0.5 (prefix-KV)
109/200
0/200
50/50
25
21
200/200
0/200
GR00T-N1.7 (backbone features)
121/200
0/200
50/50
28
19
200/200
0/200
Appendix
Table 18: Persistent intervention on Object, full pair grid. Four rollouts per condition cover all 50 valid ordered (source, target) pairs. Closer means the end effector is nearer the target than the source object. Larger shift means greater movement toward the target under injection than control. Under injection, source-task completion is 0/200 for both models.
Figure 11: ECT scene and trajectory transformations. The top row shows the original demonstration and x-axis mirror (front-back). The bottom row shows the y-axis mirror (left-right) and xy-axis mirror ( 180∘ rotation). Each panel shows four timesteps. The robot is unchanged.
Figure 12: What ECT does to conditional action support (Spatial). Each dot is a demonstration’s mean lateral action over its first ten steps, for original data (blue) or y-mirror counterfactuals (orange).
Construction
ID
Sem.
Obj.
Swap
No ECT baseline
94.2
94.8
84.4
0.0
X-mirror, paired
23.8
22.6
25.4
3.0
Y-mirror, paired
93.2
96.4
85.0
1.0
All mirrors, paired
29.0
21.8
20.2
12.0
X-mirror + X-shift, paired
97.4
99.2
95.8
1.6
XY-mirror + XY-shift, paired
95.6
98.6
91.4
7.6
Appendix
Table 19: Object-suite construction diagnostics. Success (%) for per-suite π0.5 LoRA diagnostics, separate from the main Frozen-LM comparison. X-mirror and all-mirror variants lose ID and Object competence, while Y-mirror and shifted variants preserve them.
Suite
Transform family
Action transform
Role in ECT
Spatial
mirror
reflected, then replayed
changes placement direction under fixed instruction
Object
Y-mirror / mirror+shift
reflected (x and xy also shifted), then replayed
preserves separability while changing source-object branch
Goal
mirror
reflected, then replayed
changes goal-relative spatial branch
Long
shift-augmented mix
shifted, then replayed
preserves long-horizon executability
Appendix
Table 20: ECT transform provenance. Transform choices are fixed before training by the action-validity and visual-separability criteria, not selected per evaluation rollout.
Median distance to
Minimum distance
Within ID scale
Suite
ID scale
original
transformed
to transformed
of transformed
Spatial
1.0
6.6
28.7
13.3
0%
Object
0.3
6.4
28.5
25.2
0%
Goal
0.9
11.3
24.9
19.6
0%
Long
1.7
13.9
16.2
6.4
0%
Appendix
Table 21: Layout distances from Swap states to original and transformed demonstrations. Mean planar object-and-fixture distance (cm) to the nearest layout of the same task over 500 Swap states per suite. The transformed set contains successful counterfactual layouts only. ID scale is the 90th percentile of ID-to-original distances.
Recipe
Configuration
LM
ViT
AE
Trainable
Frozen LM
_expert_only
frozen
trained
trained
844.9M
Full FT
pi05_libero (official), _full_sft (ECT)
trained
trained
trained
3,353.4M
Appendix
Table 22: Training recipes by trainable parameter group ( π0.5 , measured). Millions of parameters that receive gradients under each recipe’s trainable filter. Both recipes train the heads (2.2M).
Suite
Model
Recipe
ID
Sem.
Obj.
Swap
Task
Spatial
Official π0.5
Full FT
98.4
98.0
98.0
46.6
53.0
ECT data + ECT loss
Full FT
95.8
94.0
95.6
72.8
53.2
Standard †
Frozen LM
95.9
94.5
95.4
52.7
22.3
ECT data †
Frozen LM
96.5
93.0
95.2
72.1
18.5
ECT data + ECT loss †
Frozen LM
98.4
96.1
97.1
74.4
22.1
Object
Official π0.5
Full FT
98.6
99.0
94.4
18.2
11.0
Appendix
Table 23: Five-cell results of the four-suite π0.5 models. N=500 per cell and seed. Recipes follow Table 22 . ECT rows shaded. † Mean over three seeds (0, 1, 42), with standard deviations of ECT data + ECT loss in Table 27 .
Suite
Model
Training scope
ID
Sem.
Obj.
Swap
Task
Spatial
Official
Per suite
93.6
88.8
93.0
1.2
51.4
ECT data
Per suite
87.0
91.4
83.2
17.8
53.0
ECT data
Four suites
90.8
88.0
89.8
38.2
64.0
Object
Official
Per suite
95.6
97.6
86.0
0.0
9.0
ECT data
Per suite
97.6
98.2
88.6
0.0
10.0
ECT data
Four suites
97.2
97.2
79.8
7.0
10.2
Appendix
Table 24: GR00T-N1.7 ECT full results. Success (%) over 500 trials per cell. All ECT models use the ECT data with the standard loss. ECT rows are shaded.
ID
Swap
Task
Suite
25%
50%
100%
25%
50%
100%
25%
50%
100%
Spatial
95.2
95.8
90.0
53.6
57.4
55.4
22.6
26.8
21.2
Object
98.2
98.6
97.8
39.6
35.8
23.8
20.2
36.4
16.6
Goal
90.6
93.4
93.6
36.0
40.8
40.2
16.2
20.6
19.4
Long
94.0
90.4
94.4
14.4
15.2
15.2
16.2
24.0
15.0
Appendix
Table 25: More demonstrations of the same kind do not close the gap (prediction (vi)). The Standard model (Frozen-LM recipe, 32 examples per update) trained on 25%, 50%, and 100% of the original demonstrations, with one training seed and N=500 rollouts per cell.
Suite
Model
ID
Sem.
Obj.
Swap
Task
Spatial
Standard, then πRL
96.0
92.0
93.0
55.4
20.2
Spatial
ECT data, then πRL
95.0
91.4
94.4
69.4
28.2
Spatial
ECT data + ECT loss, then πRL
96.4
94.0
95.0
75.0
15.2
Object
Standard, then πRL
98.0
98.0
83.0
22.4
16.8
Object
ECT data, then πRL
99.2
99.4
89.0
68.6
39.0
Object
ECT data + ECT loss, then πRL
99.4
99.0
81.0
68.4
30.6
Appendix
Table 26: ECT gains survive RL post-training. Frozen-LM Standard, ECT data, and ECT data + ECT loss models after 40 PPO epochs of πRL (RLinf stock recipe, one seed). N=500 per cell. ECT rows are shaded, with the best result per column within each suite in bold.
Figure 13: Real-world training layouts and evaluation settings. The left block shows the original layout and three physically collected mirrored layouts. The right block illustrates the ID, paraphrase, attribute, unseen-position, unseen-object-and-position, and Swap conditions. The Swap comparison uses companion models trained without the layout that coincides with the swapped configuration.
Suite
Cell
Seed 0
Seed 1
Seed 42
Mean
Std.
Spatial
ID
98.2
99.0
98.0
98.4
0.5
Spatial
Sem.
97.0
96.0
95.2
96.1
0.9
Spatial
Obj.
97.2
96.6
97.6
97.1
0.5
Spatial
Swap
75.4
75.0
72.8
74.4
1.4
Spatial
Task
21.8
28.8
15.6
22.1
6.6
Object
ID
99.0
99.6
98.8
99.1
0.4
Appendix
Table 27: Seed-level success rates of π0.5 ECT data + ECT loss (Frozen LM). Std. is the sample standard deviation across three independent training seeds.
Model
Suite
Comparison
Δ Swap [95% CI]
Δ ID [95% CI]
π0.5
Spatial
ECT data + ECT loss (Frozen LM) vs. official
+27.8 [+22.7, +32.7]
0.0 [-1.3, +1.4]
π0.5
Object
ECT data + ECT loss (Frozen LM) vs. official
+52.6 [+47.3, +57.9]
+0.5 [-0.6, +1.8]
π0.5
Goal
ECT data + ECT loss (Frozen LM) vs. official
+21.3 [+16.5, +26.2]
−0.8 [-2.9, +1.4]
π0.5
Long
ECT data + ECT loss (Frozen LM) vs. official
+25.3 [+20.0, +31.3]
+1.0 [-1.7, +3.9]
GR00T-N1.7
Spatial
ECT data, per suite vs. official
+16.6 [+13.2, +20.3]
−6.6 [-10.3, -2.9]
GR00T-N1.7
Object
ECT data, per suite vs. official
0.0 [-0.8, +0.8]
+2.0 [-0.3, +4.4]
Appendix
Table 28: Key ID and Swap deltas. π0.5 intervals use a hierarchical bootstrap over seeds and rollouts and are descriptive because only three independent seeds are available. GR00T-N1.7 intervals are Newcombe hybrid-score intervals for a difference of two proportions ( N=500 each).
Scope
Retrieval-like
Correct/near-correct
Collapse
All probes
264/381 ( 69.3 [64.5, 73.7])
101/381 ( 26.5 [22.3, 31.2])
16/381 ( 4.2 [2.6, 6.7])
Goal
99/120 ( 82.5 [74.7, 88.3])
15/120 ( 12.5 [7.7, 19.6])
6/120 ( 5.0 [2.3, 10.5])
Long
63/72 ( 87.5 [77.9, 93.3])
7/72 ( 9.7 [4.8, 18.7])
2/72 ( 2.8 [0.8, 9.6])
Object
40/69 ( 58.0 [46.2, 68.9])
28/69 ( 40.6 [29.8, 52.4])
1/69 ( 1.4 [0.3, 7.8])
Spatial
62/120 ( 51.7 [42.8, 60.4])
51/120 ( 42.5 [34.0, 51.4])
7/120 ( 5.8 [2.9, 11.6])
Appendix
Table 29: Wilson confidence intervals for human rollout-label proportions. Retrieval-like includes source retrieval, other-task retrieval, and hybrid retrieval. Counts follow Table 10 .
Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow language. When presented with instructions that lack strong scene-specific supervision, VLAs suffer from counterfactual failures: they act based on vision shortcuts induced by dataset biases, repeatedly executing well-learned behaviors and selecting objects frequently seen during training regardless of language intent. To systematically study it, we introduce LIBERO-CF, the first counterfactual benchmark for VLAs that evaluates language following capability by assigning alternative instructions under visually plausible LIBERO layouts. Our evaluation reveals that counterfactual failures are prevalent yet underexplored across state-of-the-art VLAs. We propose Counterfactual Action Guidance (CAG), a simple yet effective dual-branch inference scheme that explicitly regularizes language conditioning in VLAs. CAG combines a standard VLA policy with a language-unconditioned Vision-Action (VA) module, enabling counterfactual comparison during action selection. This design reduces reliance on visual shortcuts, improves robustness on under-observed tasks, and requires neither additional demonstrations nor modifications to existing architectures or pretrained models. Extensive experiments demonstrate its plug-and-play integration across diverse VLAs and consistent improvements. For example, on LIBERO-CF, CAG improves π0.5 by 9.7% in language following accuracy and 3.6% in task success on under-observed tasks using a training-free strategy, with further gains of 15.5% and 8.5%, respectively, when paired with a VA model. In real-world evaluations, CAG reduces counterfactual failures of 9.4% and improves task success by 17.2% on average.
Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under which these failures occur, and whether they can be corrected without retraining, remain poorly understood. In this paper, we take steps toward addressing this gap. We present CorrectVLA, a framework that translates task-level natural language corrections into additive action magnitude adjustments without modifying policy weights. A human provides a single task-level correction, applied uniformly across all rollouts without per-episode intervention. In simulation, CorrectVLA recovers execution misalignment failures across both in-distribution and OOD tasks. In real-robot experiments on a UFactory xArm7 under environment shift, CorrectVLA restores near-perfect success where the base policy almost entirely breaks down, generalizing across object locations and identities. Through a taxonomy of failure modes on LIBERO-90, we find that execution misalignment failures, where the policy reaches the correct target but miscalibrates action magnitudes, represent the correctable subset, while other failure modes where semantic comprehension itself breaks down are not amenable to this approach. The approach succeeds when policies possess strategic correctness and fails when fundamental comprehension is absent, establishing a practical operational boundary for inference-time correction.
Owen Kwon, Pablo Ortega-Kral, Arthur Bucker +1
Biomedical Engineering, Carnegie Mellon University, Pittsburgh PA 15213, USA. · Robotics Institute, Carnegie Mellon University, Pittsburgh PA 15213, USA.
Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone. As a result, both human instructions and zero-shot language models can fail to reliably steer VLAs toward successful task execution. In this work, we propose a framework that interactively searches for language sequences that improve closed-loop VLA task performance, distills these sequences into a test-time language feedback policy (LFP), and learns an improvement head that predicts when language steering will improve performance. We conformalize this improvement head to prevent harmful steering interventions, where the LFP decreases task performance relative to the original instruction on out-of-distribution scenarios. Crucially, our approach operates on arbitrary frozen pre-trained VLAs, requiring neither access to the original training distribution nor fine-tuning of the underlying model. On seen environments, our conformalized LFP improves base VLA performance by 24.7% in simulation and 65.0% in hardware. On visual and semantic perturbations, our conformalized LFP has strong harmlessness guarantees, and produces recovery behaviors not observed with open-loop prompting.