Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicitly, requiring robots to recover the intended target from linguistic and perceptual context. We study this capability as Implicit Referential Grounding (IRG) and introduce RoboIRG-Bench, a manipulation benchmark designed to systematically evaluate it. Built upon RoboMME, RoboIRG-Bench contains 40 variants derived from 11 tasks and covers four challenges, including direct, reasoning-mediated, spatial, and contextual referential grounding. As IRG often requires retaining and retrieving previously established context, we evaluate representative VLAs spanning different memory mechanisms. Our evaluation reveals a noticeable referential robustness gap. Models that perform well under explicit instructions can degrade sharply when the same task-relevant information must be recovered from context. Reasoning-mediated and spatial references are particularly challenging, while models using external VLMs show greater robustness but still exhibit significant failures. Moreover, replacing the external VLM with a stronger model does not eliminate these gaps. We further validate these findings on a Franka Research 3 robot arm, where the gap persists under real-world manipulation and manifests as both incorrect referent grounding and downstream execution failures. These results establish IRG as a distinct and underexplored capability for reliable robotic instruction following and highlight the need for VLAs that can robustly integrate language, perception, reasoning, and action.
Figures & tables
Dataset
Direct
Reasoning-Mediated
Spatial
Contextual
RLBench [ 11 ]
▲
✗
✗
✗
CALVIN [ 16 ]
▲
✗
✗
✗
ARNOLD [ 9 ]
✗
✗
✗
✗
LIBERO [ 15 ]
▲
✗
✗
✗
RoboCasa [ 17 ]
▲
✗
✗
✗
VLABench [ 21 ]
✓
✗
✗
✗
TABLE I: Comparison of Implicit Referential Grounding (IRG) coverage across existing robotic manipulation benchmarks. ✓ covered; ✗ not covered; ▲ partially covered.
Fig. 2: Distribution of referring expressions over scenes, grouped by the four implicit referential grounding strategies. Each expression is counted once per scene.
Counting
Permanence
Reference
Avg
Bin Fill
Pick Xtimes
Swing Xtimes
Stop Cube
Video Unmask
Button Unmask
Video UnmaskS
Button UnmaskS
Video Repick
Video PlaceButton
Video PlaceOrder
π0.5
28.67
44.67
35.33
6.00
23.33
28.00
20.67
7.33
0.00
30.67
25.33
22.73
TTT-Expert
27.33
30.00
30.67
6.67
34.67
22.00
20.67
14.67
7.33
26.67
24.00
22.24
FrameS-Modul
40.67
90.00
93.33
50.00
33.33
28.00
24.00
15.33
30.67
56.67
42.00
45.82
GroundSG-QwenVL
55.33
88.00
4.00
0.00
90.67
23.33
26.00
15.33
29.33
42.67
32.67
37.03
Explicit
MemER
52.67
71.33
56.00
2.00
81.33
74.00
34.00
18.00
24.67
26.00
32.00
42.91
TABLE II: Success rates (%) of representative VLA models under explicit instructions and four Implicit Referential Grounding (IRG) challenges. ▼ / ▲ : decrease/increase relative to Explicit. “–” indicates not applicable.
Fig. 3: Suite-averaged success rate. Four referential strategies with implicit expressions versus the same strategy after those expressions are made explicit. Dark shows the overlap, and ∣Δ∣ is light (explicit gain) or pink (explicit drop). Δ is annotated.
Fig. 4: Representative failure modes under Implicit Referential Grounding (IRG). FramS-Modul fails to resolve referenced repetition counts under Direct Referential Grounding and target identities under Reasoning-Mediated Referential Grounding. GroundSG exhibits incorrect spatial grounding and loses track of previously introduced referents under extended context, leading to erroneous subgoal generation. Red boxes highlight failures under implicit instructions.
Model
Strategy
BinFill
PickXtimes
SwingXtimes
Avg
QwenVL
Explicit
55.33
88.00
4.00
49.11
Direct
55.33 ▲ 0.00
85.33 ▼ 2.67
8.00 ▲ 4.00
49.55 ▲ 0.44
Reasoning-Me
46.67 ▼ 8.66
66.00 ▼ 22.00
4.00 ▲ 0.00
38.89 ▼ 10.22
Contextual
43.33 ▼ 12.00
58.67 ▼ 29.33
3.33 ▼ 0.67
35.11 ▼ 14.00
Spatial
17.33 ▼ 38.00
29.33 ▼ 58.67
0.00 ▼ 4.00
15.55 ▼ 33.56
Gemini 3.1 Pro
Explicit
64.00
52.00
50.00
55.33
TABLE III: Performance Comparison between QwenVL and Gemini 3.1 Pro across three representative tasks under Explicit and four IRG settings in terms of success rate. ▼ / ▲ / ▲ : decrease/increase/no change relative to Explicit.
Fig. 5: Real-world evaluation on PutInBox and PickByCount. Green: representative successful executions under explicit instructions. Red: typical failure cases under IRG, categorized as grounding errors and execution errors.
(1)
(2)
(3)
(4)
(5)
Bin Fill
Pick Xtimes
Swing Xtimes
Avg
✓
✓
24.00
15.33
48.00
29.11
✓
✓
✓
30.67
15.33
44.67
30.22
✓
✓
✓
✓
20.00
14.67
44.67
26.45
FrameS-Modul
✓
✓
✓
✓
✓
21.33
13.33
38.67
24.44
✓
✓
41.33
81.33
8.00
43.55
✓
✓
✓
44.67
62.00
8.00
38.22
TABLE IV: Effect of increasing linguistic context on Contextual Referential Grounding. ✓ indicates a retained stage: (1) Referent Introduction, (2) Referent Reinforcement, (3) Distractor Context, (4) Action Cue, and (5) Referential Instruction.
PutInBox
PickByCount
Avg
Explicit
π0.5
90
50
70
FrameS-Modul
90
80
85
Direct
π0.5
90 ▲ 0
20 ▼ 30
55 ▼ 15
FrameS-Modul
90 ▲ 0
70 ▼ 10
80 ▼ 5
Reasoning-Me
π0.5
0 ▼ 90
10 ▼ 40
5 ▼ 65
FrameS-Modul
0 ▼ 90
20 ▼ 60
10 ▼ 75
TABLE V: Real-world task success rates (%) under explicit instructions and four IRG challenges, with each setting evaluated over 10 trials. ▼ / ▲ : decrease/no change relative to Explicit.
Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliability under Out-of-Distribution (OOD) instructions remains underexplored. In this paper, we reveal a critical failure mode in which VLA policies continue executing visually plausible actions even when the language instruction contradicts the scene. We refer to this phenomenon as linguistic blindness, where VLA policies prioritize visual priors over instruction semantics during action generation. To systematically analyze this issue, we introduce ICBench, a diagnostic benchmark constructed from the LIBERO dataset that probes language-action coupling by injecting controlled OOD instruction contradictions while keeping the visual environment unchanged. Evaluations on three representative VLA architectures, including Pi0, Pi0.5 and OpenVLA OFT, show that these models frequently succeed at tasks despite logically impossible instructions, revealing a strong visual bias in action generation. To mitigate this issue, we propose Instruction-Guided Attention Recalibration (IGAR), a train-free inference-time mechanism that rebalances attention distributions to restore the influence of language instructions. IGAR operates without retraining or architectural modification and can be directly applied to existing VLA models. Experiments across 30 LIBERO tasks demonstrate that IGAR substantially reduces erroneous execution under OOD contradictory instructions while preserving baseline task performance. We additionally validate the approach on a real Franka robotic arm, where IGAR effectively prevents manipulation triggered by inconsistent instructions.
Ninghao Zhang, Bin Zhu, Shijie Zhou +1
Tsinghua University · Singapore Management University · Institute of Trustworthy Embodied AI, Fudan University +1
Vision-Language-Action (VLA) models excel in robotic manipulation but suffer catastrophic performance drops when canonical instructions are simply paraphrased. Although this brittleness is typically addressed through costly data scaling, our probing reveals that the root cause is architectural rather than a lack of semantic understanding. Specifically, we demonstrate that current VLAs successfully retain the correct task identity internally. The failure actually stems from the joint encoding of dynamic visual observations and text, which introduces systematic feature shifts. Because the downstream action policy is highly vulnerable to these variations, it fails to translate the preserved semantics into correct control commands. To resolve this structural bottleneck, we propose Grounded Semantic Re-binding (GSR), an elegant intervention that bypasses unstable joint routing by explicitly fusing independently extracted task semantics with native visual features to train a completely re-initialized action expert from scratch. This targeted intervention dramatically restores paraphrastic invariance using only canonical demonstrations. On the LIBERO-Para benchmark, GSR improves success rates by up to 44.6 percent. It enables lightweight models to rival massively scaled baselines and pushes state-of-the-art models to a new record PRIDE score of 70.4, outperforming the recently introduced large-scale pretrained model Xiaomi-Robotics-0 in instruction generation capabilities. Building on these insights, we also introduce ParaVLA, a natively decoupled 0.33B-parameter model exhibiting near-perfect robustness to instruction rewording. Ultimately, our work proves that robust semantic grounding can be achieved through elegant structural design, bypassing the inefficient brute-force data scaling paradigm.
Zhaokai Yin, Zhipeng Zhang
1AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University · 2Research Lab, Anyverse Dynamics
Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction. Yet robot fine-tuning is optimized as imitation over task-specific action distributions, and many evaluations can be solved through visual or instruction-action shortcuts. We introduce RoboSemanticBench (RSB), an embodied benchmark for diagnosing semantic grounding in action prediction: whether post-trained VLA models can use complex instruction semantics to select and manipulate the correct physical target. In each episode, a robot receives a multiple-choice math or general-knowledge question, observes candidate answer blocks, and must grasp the block corresponding to the correct answer. RSB covers controlled arithmetic, grade-school mathematical understanding, and commonsense or factual understanding under four-choice and ten-choice suites. Across representative VLA models, we find that many policies learn to grasp candidate blocks but select the semantically correct block at near-random or below-random rates after controlling for grasp success, revealing a persistent gap between backbone-level semantic competence and action prediction.