Encoded but Not in Control: Revealing the Grounding Gap in Vision-Language Robot Policies
Organizations: HKU · CUHK(SZ) · Shanghai AI Lab · Horizon Robotics · CUHK · USTC
Abstract
Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers the task from the scene. We study this ambiguity through scene-preserving instruction interventions, using valid target substitutions, arbitrary nouns, and unrelated sentences while holding the scene fixed. We evaluate vision-language-action (VLA) policies and world-action models (WAMs) in simulation and in real-world experiments. Our analysis addresses three questions: (a) Does task success imply instruction following? When instructions request a different visible object, all evaluated policies predominantly approach and pick up the incorrect original target associated with the scene. (b) Is this failure caused by language insensitivity? Instruction perturbations affect task performance. A layerwise action lens shows intermediate action predictions respond to these perturbations. Linear probes accurately recover instructed targets, indicating modified instructions are encoded despite rarely determining target selection. (c) Why does encoded language fail to control action? Attention analysis indicates weak instruction-token contributions to action generation. Target-token attention can remain focused on the original object, revealing a mismatch between target encoding and visual grounding. UMAP and shared non-negative matrix factorization show target information remains accessible within representations increasingly organized by scene identity. Our findings expose a grounding gap concealed by nominal success and provide a diagnostic framework. They further establish a concrete criterion for progress: policies should reliably follow valid changes in user intent, even when they conflict with scene-favored behavior.
Figures & tables
| Environment | Model | Current instructed target | Original target | ||
| Tend | Success | Tend | Success | ||
| LIBERO-Object | |||||
| GR00T N1.7 | |||||
| FastWAM | |||||
| LingBot-VA | |||||
| Real World | |||||
| Benchmark | Task / Setting | Instruction | VLA | WAM | ||
| GR00T N1.7 | FastWAM | LingBot-VA | ||||
| LIBERO | Spatial | Original | ||||
| Random noun | ↑0.2 | ↓1.1 | ↓25.1 | ↓8.2 | ||
| Random sentence | ↓32.8 | ↓45.5 | ↓61.9 | ↓23.7 | ||
| Object | Original | |||||
| Random noun | ↓1.6 | ↓4.1 | ↓7.4 | ↓16.3 | ||
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Pool | Members |
| LIBERO object vocabulary used for eligibility (45) | bowl, plate, ramekin, cup, mug, box, cabinet, drawer, basket, book, butter, cheese, chocolate, cream, sauce, can, bottle, pan, pot, lid, tray, rack, shelf, table, desk, chair, cloth, towel, bag, container, toy, block, cube, ball, hammer, brush, fork, spoon, knife, glass, jar, kettle, pitcher, bin, bucket. |
| Replacement and random-sentence nouns (99) | cat, dog, bird, fish, horse, lion, tiger, bear, wolf, rabbit, deer, eagle, shark, whale, frog, snake, turtle, rock, stone, tree, leaf, cloud, river, mountain, beach, forest, lake, desert, island, volcano, cave, valley, lamp, clock, mirror, pillow, blanket, carpet, curtain, vase, candle, picture, statue, phone, remote, cable, apple, banana, orange, grape, pizza, bread, rice, soup, cookie, cake, sandwich, salad, pasta, egg, milk, wrench, nail, screw, drill, rope, chain, key, lock, pipe, wire, tape, glue, paint, wheel, gear, shirt, hat, shoe, sock, coat, glove, belt, scarf, coin, stamp, ticket, card, map, flag, sign, label, button, switch, ring, medal, trophy, balloon, umbrella. |
| Random-sentence verbs (10) | observe, locate, inspect, examine, survey, approach, avoid, ignore, study, scan. |
| Random-sentence adjectives (10) | curious, ancient, shiny, purple, enormous, tiny, mysterious, hollow, rusty, invisible. |
| No. | Template |
| 1 | The {adj} {noun1} is resting next to the {noun2}. |
| 2 | A {noun1} and a {noun2} are placed on the {noun3}. |
| 3 | Please {verb} the {adj} {noun1} carefully. |
| 4 | There is a {noun1} near the {noun2} and a {noun3} on the floor. |
| 5 | The robot should {verb} the {noun1} without touching the {noun2}. |
| 6 | A {adj} {noun1} has been spotted beside the {noun2}. |
| Scene | Target set | Scene | Target set | ||
| 1 | alphabet soup , salad dressing , cream cheese , milk , tomato sauce , butter | 2 | cream cheese , alphabet soup , milk , tomato sauce , butter , orange juice | ||
| 3 | milk , cream cheese , tomato sauce , butter , orange juice , chocolate pudding | 4 | tomato sauce , milk , butter , orange juice , chocolate pudding , bbq sauce | ||
| 5 | butter , tomato sauce , orange juice , chocolate pudding , bbq sauce , ketchup | 6 | orange juice , butter , chocolate pudding , bbq sauce , ketchup , salad dressing | ||
| 7 | chocolate pudding , orange juice , bbq sauce , ketchup , salad dressing , alphabet soup | 8 | bbq sauce , chocolate pudding , ketchup , salad dressing , alphabet soup , cream cheese | ||
| 9 | ketchup , bbq sauce , salad dressing , alphabet soup , cream cheese , milk | 10 | salad dressing , ketchup , alphabet soup , cream cheese , milk , tomato sauce |
| Model | Denoising steps | Layers / step | Evaluated action region | Instantiation of |
| 10 | 18 | Timestep-conditioned adaptive normalization followed by the native action_out_proj | ||
| GR00T N1.7 | 4 | 32 | Native output normalization and timestep-conditioned modulation, followed by the output projection and LIBERO-specific embodiment mapping | |
| FastWAM | 10 | 30 | Shared trained ActionDiT action head used in the native inference pathway | |
| LingBot-VA | 10 | 30 | * | Timestep-conditioned norm_out followed by action_proj_out |