Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot serve as direct targets because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones; it also improves compositional generalization on a real robot. Project website: https://taegeunyang.github.io/craft/
Figures & tables
Figure 1: Generalization to undemonstrated skill combinations. Fine-tuned VLAs struggle to execute new combinations of demonstrated skills. CRAFT uses demonstrated skill executions to train the policy to follow instructions for new combinations, improving success on undemonstrated combinations across two benchmarks and three VLA models over standard fine-tuning.
Figure 2: Overview of CRAFT. Yellow instruction highlights mark the currently required skill. (1) We encourage reuse of skill representations across executions by contrasting prediction errors while holding the state representation and target velocity fixed (target velocity and different-skill replacements omitted for clarity). (2) We combine the counterfactual skill representation with a state representation from a demonstration of the required skill. The demonstration’s target velocity supervises the prediction, training the counterfactual skill representation to reflect the required skill (skill-changing case shown).
Figure 3: Compositional benchmarks with example demonstrations and task-combination matrices (Appendix C ). Demonstrations cover all constituent skills but only a subset of combinations (checked cells); empty cells denote undemonstrated combinations. Rows and columns indicate cube and plate colors, respectively. The two Pick-Place-Press matrices correspond to red and blue buttons, respectively.
π0.5
π0
GR00T N1.7
Method
Demo.
Undem.
Demo.
Undem.
Demo.
Undem.
Pick–Place
FT ( Full )
100.0±0.0
8.4±0.3
100.0±0.0
2.8±0.2
98.8±0.3
1.1±0.1
FT ( LoRA )
100.0±0.0
46.3±0.9
97.5±0.0
1.0±0.2
97.8±0.6
2.6±0.3
FT ( Frozen )
99.0±0.0
21.1±0.2
48.2±1.4
3.0±0.5
92.3±1.3
2.2±0.1
CAG-VA a
96.3±1.5
66.5±0.5
50.5±1.3
6.8±0.6
88.7±1.3
4.9±0.5
Table 1: Success rates (%) on demonstrated (Demo.) and undemonstrated (Undem.) combinations across three VLA models and two benchmarks. We report means and standard deviations over three evaluation seeds.
Figure 4: Responses to instruction changes. Yellow highlights identify the entity assigned to the current operation. The robot approaches the red cube and maintains this behavior when only the future placement target changes (1 → 2 → 3). It redirects its approach when the picking target changes (3 → 4), then changes its destination during placement (5 → 6), placing the blue cube on the yellow plate.
Figure 5: Skill representations learned with and without Lskill on Pick-Place , independently projected with t-SNE. Each point represents a demonstration frame, colored by the entity of the currently required skill. Shaded regions indicate pick and place groups.
Auxiliary objectives
π0.5
GR00T N1.7
Skill
Pres.
Change
Demo.
Undem.
Demo.
Undem.
88.8±0.6
1.9±0.4
99.3±0.3
3.8±0.1
✓
99.2±0.6
0.3±0.4
97.2±1.0
6.1±0.5
✓
✓
99.8±0.3
0.4±0.2
97.3±0.6
5.3±0.3
✓
✓
97.7±0.3
14.7±0.8
94.5±1.3
12.1±0.4
CRAFT (all)
99.0±0.9
84.2±0.3
99.0±0.5
72.1±0.9
Table 2: Success rates (%) on Pick-Place for training objective ablations. (mean ± std.; three evaluation seeds)
Figure 6: Real-robot execution of an undemonstrated combination. The instruction requires picking the blue cube and placing it on the green plate; this combination is absent from the demonstrations, although both constituent skills have been demonstrated. Both CRAFT and Full pick the instructed blue cube, but CRAFT places it on the instructed green plate, whereas Full places it on the blue plate, corresponding to the demonstrated blue-to-blue combination.
Method
Demo.
Undem.
FT ( Full )
19/20
9/60
CRAFT (ours)
19/20
43/60
Table 3: Real-robot task success (successful trials / total trials) on demonstrated (Demo.) and undemonstrated (Undem.) combinations.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1: An example initial scene and the six pick-and-place tasks from the LIBERO- Goal suite. The first panel shows the initial scene, and the remaining panels show the six tasks with their instructions.
cabinet top
rack
plate
stove
bowl
wine bottle
✓ (1)
✓ (4)
∘
∘
bowl
✓ (2)
✓ (3)
✓ (6)
cream cheese
∘
∘
∘
✓ (5)
Appendix
Table A.1: Object–target combinations derived from the six pick-and-place tasks in LIBERO- Goal . ✓ marks a demonstrated combination, with the corresponding task index from Table A.2 ; ∘ marks an undemonstrated combination used for evaluation. Blank cells denote undemonstrated combinations not evaluated.
Task
π0
π0.5
π0 -FAST
GR00T N1.7
pick-and-place
1
put the wine bottle on top of the cabinet
98.0±1.6
98.0±1.6
97.3±0.9
98.7±0.9
2
put the bowl on top of the cabinet
97.3±1.9
96.7±0.9
93.3±6.6
94.7±1.9
3
put the bowl on the plate
100.0±0.0
99.3±0.9
98.0±0.0
96.0±1.6
4
put the wine bottle on the rack
75.3±1.9
95.3±2.5
46.0±1.6
96.0±1.6
5
put the cream cheese in the bowl
98.7±0.9
98.0±0.0
99.3±0.9
93.3±0.9
Appendix
Table A.2: Success rate (%) on the ten original LIBERO- Goal tasks. Released checkpoints are evaluated on 50 trials per task and evaluation seed; we report mean ± std over three seeds. Task indices are used for cross-reference with Tables A.1 and A.3 .
Instruction
π0
π0.5
π0 -FAST
GR00T N1.7
put the wine bottle on the plate
6.0±0.0
52.0±3.3
0.0±0.0
0.0±0.0
put the wine bottle on the stove
2.0±1.6
77.3±4.1
0.7±0.9
0.0±0.0
put the cream cheese on top of the cabinet
0.0±0.0
78.7±2.5
7.3±1.9
6.7±2.5
put the cream cheese on the plate
4.0±1.6
24.7±7.7
88.0±1.6
0.7±0.9
put the cream cheese on the stove
2.7±0.9
22.0±2.8
1.3±0.9
2.7±1.9
mean
2.9±0.4
50.9±2.4
19.5±0.7
2.0±1.0
Appendix
Table A.3: Success rate (%) on the five constructed undemonstrated combinations. We report mean ± std over three evaluation seeds, with 50 trials per task and seed.
Figure A.2: Failure on an undemonstrated combination in LIBERO- Goal . Given “put the wine bottle on the stove,” the fine-tuned π0.5 policy grasps the instructed object but places it on the rack, corresponding to the demonstrated “put the wine bottle on the rack” task.
Model
Instructed task
Shares the object
Shares the target
Other of the ten
None
π0
2.9±0.4
9.6±0.7
48.8±0.6
0.5±0.4
38.1±1.2
π0.5
50.9±2.4
38.4±1.2
2.8±0.3
0.0±0.0
7.9±0.9
π0 -FAST
19.5±0.7
41.9±2.0
4.1±0.2
0.8±0.7
33.7±0.8
GR00T N1.7
2.0±1.0
39.3±1.0
25.5±1.6
1.9±0.7
31.3±1.0
Appendix
Table A.4: Outcomes of rollouts on the five constructed undemonstrated combinations. Values are percentages of trials pooled across combinations, reported as mean ± std over three seeds. None indicates that the rollout satisfies neither the instructed task goal nor any of the ten original LIBERO- Goal task predicates.
Figure A.3: Behavior of the fine-tuned π0.5 policy under an empty instruction. Without a task instruction, the policy turns on the stove in one trial (top) and puts the bowl on the plate in another (bottom), both corresponding to demonstrated tasks.
Figure B.1: CRAFT architecture and attention masks. (a) Original and CRAFT VLM–action expert interfaces. CRAFT conditions the π -family action expert on layer-wise K/V from skill and state queries, and GR00T on their final hidden states. (b) Attention masks used to construct the two representations. Rows attend to columns; filled cells permit attention, white cells block it, and triangular cells denote causal attention. The π family extracts the two query representations in separate shared-backbone passes, whereas GR00T extracts both in a single pass. Token counts are schematic; formatting and padding tokens are omitted.
Variant
Pick (completed)
Place (current)
Press (future)
Original
Blue cube
Blue plate
Blue button
Skill-changing
Blue cube
Red plate
Blue button
Skill-preserving
Blue cube
Blue plate
Red button
Appendix
Table B.1: Counterfactual instructions during placing in the pick–place–press task, whose demonstrations contain only all-blue and all-red combinations. Bold entries indicate changed entity assignments. Both counterfactuals retain the completed pick assignment and specify an undemonstrated combination.
Loss
Skill representation
State representation
Action expert
Lfm
Gradient passes
Gradient passes
Updated
Lskill
Gradient passes
Gradient stopped
Updated
Lcf
Gradient passes
Gradient stopped
Fixed for this loss
Appendix
Table B.2: Gradient paths for the training objectives. Each row describes the contribution of one loss to the joint update.
Figure C.1: Benchmark construction for (a) Pick-Place and (b) Pick-Place-Press . Top: demonstrated same-color combinations used for data collection. Bottom: example initial scenes from each benchmark. The remaining combinations form the undemonstrated evaluation set.
Figure D.1: CRAFT success rates (%) across Pick-Place task combinations, averaged over three evaluation seeds. Rows and columns specify cube and plate colors, respectively. Hatched cells denote demonstrated combinations; all panels use the same color scale.
Figure D.2: CRAFT success rates (%) across Pick-Place-Press task combinations, averaged over three evaluation seeds. Rows and columns within each matrix specify cube and plate colors, respectively; the upper and lower matrices correspond to red and blue buttons. Hatched cells denote demonstrated combinations; all panels use the same color scale.
Figure D.3: Action predictions under instruction changes with a fixed observation. Top: the observation is taken while picking the blue cube, and the instructed cube is varied across four colors. Bottom: the observation is taken while placing on the red plate, and the instructed plate is varied across four colors. The blue target in the top row and red target in the bottom row correspond to the original executions. Trajectory colors indicate the instructed entity, with predicted action chunks shown for 30 random seeds. Insets enlarge the marked regions at the same scale within each row.
Figure D.4: Skill-representation reuse across executions on Pick-Place with π0.5 . For each recipient observation, skill-query K/V are replaced with those from another execution while the recipient’s state representation and sampling noise remain fixed. Rows and columns indicate recipient and donor entities, respectively. Each cell reports the mean relative change (%) in the predicted action chunk. Outlined diagonal cells use donors executing the same skill as the recipient; off-diagonal cells use donors executing the same operation with a different entity. Both panels use the same color scale.
Figure D.5: Query-count and allocation ablations on Pick-Place . (a) Total query count for π0.5 and GR00T N1.7, with an equal number of skill and state queries. (b) Skill/state query allocation for π0.5 with 64 total queries. Solid and dashed lines denote success on undemonstrated and demonstrated combinations, respectively.
Figure D.6: CAG-VA guidance-strength sweeps on Pick-Place and Pick-Place-Press . Each configuration uses a separately fine-tuned vision-only VA policy together with the corresponding language-conditioned VLA policy. Curves show success on undemonstrated combinations as a function of guidance strength ω . The initial sweep uses seed 7; selected settings additionally evaluated with seeds 0 and 42 are shown as three-seed means with standard-deviation error bars, while the remaining points show seed-7 results. Outlined markers denote the best guided setting ( ω>1 ) in each seed-7 sweep; ω=1 recovers the unguided language-conditioned policy.
Configuration
ω=1
1.5
2
3
4
5
Pick-Place
π0.5 LoRA
46.3(100)
55.8(100)
60.5(100)
64.3(100)
65.7(99.5)
66.5±0.5(96.3±1.5)
π0.5 Full
8.4(100)
11.0(100)
13.7(99.5)
16.2(98.5)
15.2(95)
15.5(94.5)
π0.5 Frozen
21.1(99)
22.7(98.7)
21.8(97.5)
19.8(95.5)
17.8(86)
15.7(83)
π0 LoRA
1.0(97.5)
2.5(96.5)
3.8(93)
5.0(81)
4.0(71)
2.7(57)
π0 Full
2.8(100)
4.0(100)
4.0(100)
4.0(97.5)
3.7(93)
4.2(87)
Appendix
Table D.1: CAG-VA guidance-strength sweep. Each cell reports undemonstrated success with demonstrated success in parentheses (%). The initial sweep uses seed 7; selected settings additionally evaluated with seeds 0 and 42 are reported using their three-seed means. Bold entries indicate the final CAG-VA configurations reported in Table 1 and include mean ± standard deviation over the three evaluation seeds. The setting ω=1 corresponds to the unguided language-conditioned policy.
Figure D.7: Additional real-robot rollouts of CRAFT on undemonstrated Pick-Place combinations. Each row shows a different instruction and its execution over time. CRAFT picks the instructed cube and places it on the instructed plate, although the corresponding cube–plate combination is absent from the demonstrations.