cs.CVOct 4, 2026

GeoBridge-VLA: Geometry-Aware Residual Adaptation for Vision-Language-Action Models

Authors: Hyun Song, Kangmin Kim, Loren Jinsoo Um, Minhui Han, Jaehyeok Park, Taewan Cho, Andrew Jaeyong Choi

Organizations: School of Computing, iRASC Lab., Gachon University

Abstract

Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's frozen visual encoder and using them for action prediction. Stage I trains a feature bridge and geometry decoder with depth supervision. Stage II freezes these modules and trains a gated residual interface together with the action-side projections and action expert. The residual augments the existing visual tokens without adding a second image encoder or increasing the token count. Deployment requires RGB, robot state, and language, but no depth observations. Under matched evaluation conditions, GeoBridge-VLA achieves 70.9% success on LIBERO, compared with 60.0% for SmolVLA. Disabling the residual in the same trained checkpoint reduces success from 70.90% to 69.85%, with mixed effects across suites. On a physical ROBOTIS OMY robot, GeoBridge-VLA succeeds in 148 of 200 trials (74.0%) across four tasks, compared with 108 of 200 (54.0%) for SmolVLA.

Figures & tables

Explore similar work

CardsList
  1. G3^3VLA: Geometric inductive bias for Vision-Language-Action Models

    Jun 23, 2026Yue Peng, Yongzhe Zhao, Artur Habuda +5Diffusion-Based Vision-Language-ActionsVisual Geometry Grounded Transformer

  2. Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference

    May 4, 2026Yudong Liu, Yuan Li, Zijia Tang +12Robotic ManipulationLatent Variable