Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance
Organizations: Case Western Reserve University Cleveland, OH, USA
Abstract
While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visual shortcuts, associating actions with task-irrelevant visual features rather than the intended task semantics. These shortcuts block recomposition of elements already seen by the policy, that is, compositional generalization. Existing approaches mitigate such entanglement through task-relevant perception or targeted data diversification, but offer no explicit mechanism for unseen recomposition and require backbone-specific modifications with retraining. We observe that under such recomposition, VLAs often fail at global grounding while retaining local manipulation skills that recover near the correct target in familiar configurations. Therefore, we propose Referential Guidance (ReGuide), a training-free wrapper that, given object poses from a grounding module, combines semantic and geometric rebinding to guide the end-effector into demonstration-supported configurations of the instructed referent, where the frozen policy can resume execution. Experiments in simulation across multiple VLA backbones as well as on a real robot show that ReGuide improves success rates under compositional shifts by up to 56.8 and 75.0 percentage points, respectively, while preserving standard-task performance.
Figures & tables
| Method | Year | Orig. | Recombination Axes | Average | |||||
|---|---|---|---|---|---|---|---|---|---|
| Target Region | Target Object | Unseen Object | Task Scene | All | Comp. Avg. | ||||
| Found. | OpenVLA-OFT | 2025 | |||||||
| 2025 | |||||||||
| 2025 | |||||||||
| X-VLA | 2026 | ||||||||
| GR00T N1.7 | 2026 | ||||||||
| Method | Pick-and-Place | Push | Average (%) | |||||
|---|---|---|---|---|---|---|---|---|
| Orig. | Comp. | Unseen | Orig. | Comp. | Orig. | Comp. | All | |
| Bare | 26/30 | 7/60 | 8/20 | 17/20 | 0/20 | |||
| + ReGuide | 28/30 | 53/60 | 17/20 | 18/20 | 20/20 | |||
| Variant | Recombination Axes | Comp. Avg. | vs Full | |||
|---|---|---|---|---|---|---|
| Target Region | Target Object | Unseen Object | Task Scene | |||
| Bare | ||||||
| Naive transport | ||||||
| w/o competing referents | ||||||
| w/o chunk commitment | ||||||
| w/o pre-contact sets | ||||||
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| , , , | frozen policy, action chunk, chunk length, executed actions per replan |
| , | translation, rotation and gripper command, end-effector pose |
| , | command scales of the controller, tracking gains fitted on demonstrations |
| , , , | scene entity, its position, orientation and bounding-box half-extents |
| , | instructed referent of stage , moving point (end-effector or held object) |
| , , | unexecuted chunk remainder, predicted endpoint and posture |
| Statistic | Measured on | Level and role |
|---|---|---|
| tracking gains , | realized against commanded motion, per demonstration | , unit conversion |
| step bounds , | per-step displacement and rotation, approach and carry | , upper bound |
| commitment radius | distance to the box at the entry event, per object type | , trigger |
| carry radius | distance to the destination during demonstrated carries, per class | , trigger |
| entry altitude | height above the object top at the entry event, per type | , target |
| grasp offset , rotation | end-effector pose at engagement in the canonical frame, per slice | median and medoid, target |
| Cell | Host scene and anchor task | Change | Instruction |
|---|---|---|---|
| A1 | Spatial, bowl at table centre to plate | destination is the stove | pick up the black bowl from table center and place it on the stove |
| A2 | Goal, cream cheese to bowl | destination is the open top drawer | put the cream cheese in the top drawer of the wooden cabinet |
| A3 | Goal, bowl to plate | destination is the open top drawer | put the black bowl in the top drawer of the wooden cabinet |
| A4 | Goal, wine bottle to rack | destination is the stove | put the wine bottle on the stove |
| B1 | Living room 6, white mug to plate | the mug is replaced in place by a moka pot | put the moka pot on the plate |
| B2 | Living room 6, white mug to plate | the target is the chocolate pudding of the scene | put the chocolate pudding on the plate |
| Method | Checkpoint | Execution |
|---|---|---|
| , | official LIBERO checkpoints of openpi | official client, 5 actions executed per replan, chunks of 50 actions for and 10 for |
| OpenVLA-OFT | official checkpoint trained on the four suites | chunk of 8 executed open loop, two images and proprioception |
| X-VLA | official LIBERO checkpoint | official client with absolute end-effector targets, the full chunk of 30 actions executed open loop |
| GR00T N1.7 | fine-tuned by us on the four suites, see below | official horizon, 16 actions predicted and 8 executed |
| OTTER | trained by us with the official code, as no checkpoint is released, see below | official evaluation loop, one action per step with a context of 12 frames |
| GuidedVLA | released checkpoint, built on , see below | official server with corrected depth-adapter checkpoint loading |
| Metric | Task type | Default start | Pre-contact start |
|---|---|---|---|
| Grasp | Original | 400/400 | 362/400 |
| Recomposed | 258/400 | 383/400 | |
| Task success | Original | 397/400 | 343/400 |
| Recomposed | 136/400 | 266/400 |
| Method | A1 | A2 | A3 | A4 | B1 | B2 | B3 | B4 | C1 | C2 | C3 | C4 | D1 | D2 | D3 | D4 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenVLA-OFT | 0 | 0 | 27 | 0 | 3 | 0 | 0 | 0 | 3 | 0 | 31 | 35 | 0 | 13 | 0 | 0 | 112 |
| 0 | 0 | 10 | 0 | 0 | 0 | 3 | 0 | 0 | 7 | 0 | 47 | 0 | 0 | 1 | 0 | 68 | |
| 1 | 5 | 42 | 33 | 4 | 13 | 9 | 24 | 3 | 19 | 27 | 50 | 0 | 42 | 37 | 3 | 312 | |
| X-VLA | 0 | 47 | 31 | 0 | 1 | 6 | 12 | 0 | 0 | 2 | 16 | 50 | 0 | 0 | 0 | 0 | 165 |
| GR00T N1.7 | 0 | 0 | 8 | 21 | 2 | 1 | 10 | 6 | 0 | 1 | 9 | 46 | 0 | 0 | 0 | 0 | 104 |
| OTTER | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 15 | 11 | 0 | 0 | 1 | 0 | 27 |
| Method | O1 | O2 | O3 | O4 | O5 | O6 | O7 | O8 | O9 | O10 | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenVLA-OFT | 49 | 49 | 50 | 50 | 46 | 46 | 50 | 50 | 48 | 50 | 488 |
| 49 | 49 | 40 | 50 | 5 | 40 | 49 | 48 | 50 | 50 | 430 | |
| 50 | 50 | 47 | 50 | 31 | 44 | 50 | 50 | 48 | 50 | 470 | |
| X-VLA | 49 | 49 | 46 | 50 | 48 | 47 | 50 | 50 | 50 | 49 | 488 |
| GR00T N1.7 | 50 | 49 | 49 | 49 | 34 | 44 | 50 | 50 | 48 | 50 | 473 |
| OTTER | 46 | 42 | 31 | 50 | 21 | 30 | 41 | 50 | 50 | 50 | 411 |
| Variant | A1 | A2 | A3 | A4 | B1 | B2 | B3 | B4 | C1 | C2 | C3 | C4 | D1 | D2 | D3 | D4 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Naive transport | 35 | 2 | 0 | 50 | 7 | 50 | 50 | 34 | 16 | 38 | 48 | 49 | 39 | 34 | 46 | 50 | 548 |
| w/o competing referents | 34 | 49 | 45 | 45 | 49 | 49 | 14 | 42 | 44 | 48 | 44 | 47 | 38 | 43 | 48 | 48 | 687 |
| w/o chunk commitment | 49 | 41 | 43 | 45 | 45 | 49 | 50 | 23 | 25 | 47 | 49 | 42 | 44 | 44 | 49 | 46 | 691 |
| w/o pre-contact sets | 50 | 48 | 43 | 33 | 28 | 49 | 49 | 15 | 23 | 29 | 50 | 49 | 47 | 40 | 29 | 25 | 607 |
| w/o bounded transport | 50 | 31 | 41 | 46 | 45 | 41 | 49 | 40 | 49 | 50 | 38 | 43 | 41 | 44 | 47 | 46 | 701 |
| ReGuide (full) | 50 | 48 | 46 | 48 | 49 | 49 | 49 | 39 | 48 | 46 | 46 | 49 | 49 | 42 | 48 | 50 | 756 |
| Quantity | Default | Variant | Success (%) | Default only | Variant only | ||
| Default | – | – | 756 | 94.5 ±1.6 | – | – | – |
| 3 | 1 | 760 | 95.0 ±1.5 | 27 | 31 | 0.694 | |
| 3 | 5 | 759 | 94.9 ±1.5 | 25 | 28 | 0.784 | |
| (mm) | 4 | 2 | 755 | 94.4 ±1.6 | 34 | 33 | 1.000 |
| (mm) | 4 | 8 | 752 | 94.0 ±1.7 | 36 | 32 | 0.716 |
| trigger level | 766 | 95.8 ±1.4 | 25 | 35 | 0.245 |
| Cell | Full | No assistance |
|---|---|---|
| A1 | 50 | 48 |
| A2 | 48 | 49 |
| A3 | 46 | 43 |
| A4 | 48 | 48 |
| B1 | 49 | 35 |
| B2 | 49 | 48 |
| Grounding | Target Region | Target Object | Unseen Object | Task Scene | Comp. Avg. |
|---|---|---|---|---|---|
| Oracle | |||||
| Perception |
| Action | No hand-back | Before grasp | After grasp |
|---|---|---|---|
| VLA (full) | 120/139 (86.3) | 510/531 (96.0) | 126/130 (96.9) |
| Default | 126/147 (85.7) | 52/513 (10.1) | 127/140 (90.7) |
| Hold | 121/141 (85.8) | 0/529 (0.0) | 103/130 (79.2) |
| Cell | Full | Default | Hold | Cell | Full | Default | Hold |
|---|---|---|---|---|---|---|---|
| A1 | 50 | 50 | 49 | C1 | 48 | 0 | 0 |
| A2 | 48 | 2 | 1 | C2 | 46 | 1 | 0 |
| A3 | 46 | 9 | 9 | C3 | 46 | 0 | 0 |
| A4 | 48 | 47 | 42 | C4 | 49 | 0 | 0 |
| B1 | 49 | 0 | 0 | D1 | 49 | 34 | 0 |
| B2 | 49 | 15 | 13 | D2 | 42 | 42 | 37 |
| Method | A | B | C | D | Comp. | Orig. | All |
|---|---|---|---|---|---|---|---|
| Bare | 1/4 | 0/4 | 2/4 | 1/4 | 4/16 | 10/10 | 14/26 |
| Harness VLA | 3/4 | 1/4 | 2/4 | 4/4 | 10/16 | 8/10 | 18/26 |
| ReGuide | 4/4 | 4/4 | 3/4 | 4/4 | 15/16 | 8/10 | 23/26 |
| Measurement | Bare | ReGuide |
|---|---|---|
| Policy query | 127.2 / 128.1 | 126.7 / 127.7 |
| Replanning latency | 129.0 / 129.9 | 129.2 / 130.3 |
| Client computation per step | 1.60 / 1.81 | 2.14 / 2.61 |
| Client computation (99th percentile) | 2.04 | 3.49 |
| Simulator step (median) | 10.3 | 10.1 |