Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's frozen visual encoder and using them for action prediction. Stage I trains a feature bridge and geometry decoder with depth supervision. Stage II freezes these modules and trains a gated residual interface together with the action-side projections and action expert. The residual augments the existing visual tokens without adding a second image encoder or increasing the token count. Deployment requires RGB, robot state, and language, but no depth observations. Under matched evaluation conditions, GeoBridge-VLA achieves 70.9% success on LIBERO, compared with 60.0% for SmolVLA. Disabling the residual in the same trained checkpoint reduces success from 70.90% to 69.85%, with mixed effects across suites. On a physical ROBOTIS OMY robot, GeoBridge-VLA succeeds in 148 of 200 trials (74.0%) across four tasks, compared with 108 of 200 (54.0%) for SmolVLA.
Figures & tables
Fig. 2: Geometry integration in VLA policies. (a) A native RGB-only VLA uses its visual tokens for action conditioning. (b) A separate-depth design adds a second RGB encoder and a fusion module. (c) GeoBridge-VLA shares the native visual encoder, decodes geometry through a feature bridge and K4-based geometry decoder, and adds a learned residual to the existing visual tokens. Language and robot-state inputs are omitted here for clarity.
Fig. 3: The four LIBERO evaluation suites [ 31 ] . Representative RGB observations and language instructions illustrate spatial relations (Spatial), object variation (Object), goal variation (Goal), and multi-step manipulation (Long). These panels introduce the task settings used in Table II .
Fig. 4: ROBOTIS OMY task settings. Each task panel pairs the wrist camera (left) with the external camera (right). Top row: T1, red-cube pick-and-place; T2, drawer closing. Bottom row: T3, bowl retrieval and transfer; T4, size-ordered block stacking. The photos document camera views and manipulation settings. Table V reports task success rates for SmolVLA and GeoBridge-VLA.
Task
Language instruction
T1: Cube transfer
Pick up the red cube on the right side and place it on the blue plate on the left side.
T2: Drawer closing
Close the drawer.
T3: Bowl transfer
Pick up the red bowl in the drawer and place it on the blue plate on the right side.
T4: Block stacking
Stack the three red cubes on the blue plate, from largest at the bottom to smallest at the top.
TABLE I: OMY tasks and language instructions.
LIBERO Suite
SmolVLA (SR)
GeoBridge-VLA (SR)
Change ( Δ , pp)
Spatial
68.8%
73.4%
+4.6
Object
65.6%
86.4%
+20.8
Goal
72.8%
76.6%
+3.8
Long
32.8%
47.2%
+14.4
Overall
60.0%
70.9%
+10.9
TABLE II: LIBERO success rates (%). Both methods use the same evaluation conditions. Overall is the suite average; Δ is the change in percentage points.
Fig. 5: LIBERO-Long task sequence. Frames 0, 125, and 251 illustrate progress through a multi-step mug-and-pudding manipulation task.
Model
AbsRel ↓
RMSE (m) ↓
δ1↑
SmolVLA shared encoder
0.030
0.1316
0.975
DA3-Mono-Large [ 13 ]
0.036
0.1549
0.988
TABLE III: Shared-encoder depth prediction. Results on the LIBERO front-view validation set.
Fig. 6: Shared-encoder depth predictions. Qualitatively selected LIBERO front views show RGB (left), reference depth (middle), and predictions (right), illustrating robot, object, and tabletop geometry.
Fig. 7: OMY shared-encoder depth predictions. RGB observations (left) and predicted relative depth (right) for (a) the front view and (b) the wrist view. Examples are qualitatively selected from the training set.
LIBERO Suite
Normal
Zero
Δ (pp)
Spatial
73.4%
70.2%
+3.20
Object
86.4%
87.6%
-1.20
Goal
76.6%
77.2%
-0.60
Long
47.2%
44.4%
+2.80
Overall
70.9%
69.85%
+1.05
TABLE IV: Same-checkpoint residual intervention (%). Normal and zero use the same trained checkpoint without retraining. Zero sets Δz=0 at inference; Δ is normal minus zero in percentage points; positive values indicate higher success with the residual enabled.
Task
SmolVLA
GeoBridge-VLA
Δ (pp)
T1: cube transfer
60.0%
80.0%
+20.0
T2: drawer closing
70.0%
86.0%
+16.0
T3: bowl transfer
50.0%
70.0%
+20.0
T4: block stacking
36.0%
60.0%
+24.0
Overall
54.0%
74.0%
+20.0
TABLE V: OMY success rates (%). Each method is evaluated over 50 trials per task. Overall is the unweighted mean across the four tasks; Δ denotes GeoBridge-VLA minus SmolVLA in percentage points.
Decoder
AbsRel ↓
RMSE (m) ↓
δ1↑
SPACE-CLIP V1
0.05239
0.06134
0.95497
SPACE-CLIPv2
0.04321
0.04993
0.96678
TABLE VI: Standalone depth comparison. Matched CLIP backbone and 9,984-image LIBERO test split.