Generalized referring expression segmentation (GRES) requires dynamically balancing high-level semantics for identifying a variable number of language-specified referents with fine-grained visual evidence for precise boundary delineation. This requirement challenges existing cascaded vision-language architectures, which typically rely on static feature interfaces and single-pass mask prediction, limiting adaptive perception and geometric correction. We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks. Evolution-Aware Semantic Scheduling (EASS) couples contour-guided bidirectional boundary sampling with state-conditioned routing of multilevel multimodal features, adapting perception to each contour state. Following supervised initialization, Dustbin-Augmented Entropic Credit Transport GRPO (DECT-GRPO) jointly optimizes discrete grounding and continuous contour actions with instance-level credits. Its rollout rewards and credits are derived from soft prediction-target correspondences that account for false positives and missed targets. ContourVLA improves gIoU over the strongest evaluated baselines by 8.7, 2.8, and 2.7 points on gRefCOCO val, testA, and testB, respectively, and achieves the highest mIoU across all eight RefCOCO, RefCOCO+, and RefCOCOg splits.
Figures & tables
Figure 1: Paradigm-level comparison of C o n t o u r V L A with prior methods.
Figure 2: Overview of the C o n t o u r V L A framework. VLIM predicts grounding boxes for contour initialization and provides visual and multilevel multimodal features. EASS performs adaptive bidirectional boundary sampling and routes semantic representations according to the evolving state. GAD combines contour point tokens with instance context and routed semantics to predict pointwise geometric action chunks. Executing these actions updates the contours and conditions subsequent observations, forming a closed perception-action loop. The dashed inset illustrates a GAD block.
Figure 3: Overview of DECT-GRPO. Dustbin-augmented entropic transport establishes soft prediction–target correspondences and assigns instance-level credits to grounding tokens and contour actions for joint policy optimization.
Method
gRefCOCO
RefCOCO
RefCOCO+
RefCOCOg
val
testA
testB
val
testA
testB
val
testA
testB
val
test
gIoU
cIoU
gIoU
cIoU
gIoU
cIoU
mIoU
cIoU
mIoU
cIoU
mIoU
cIoU
mIoU
cIoU
mIoU
cIoU
mIoU
cIoU
mIoU
cIoU
mIoU
cIoU
Grounded-SAM ( Ren et al., 2024 )
46.0
44.1
50.8
48.7
43.5
41.3
46.2
44.4
48.6
47.0
42.9
40.9
38.2
36.5
42.4
40.5
33.0
30.8
44.1
42.0
44.5
43.0
LISA-7B ( Lai et al., 2024 )
61.4
61.7
66.0
68.5
58.8
60.6
68.1
74.6
70.6
79.3
64.5
72.3
54.8
65.1
60.3
70.8
48.4
58.1
63.1
67.9
66.2
70.6
LISA-13B ( Lai et al., 2024 )
63.4
62.9
68.1
69.6
61.8
62.2
75.3
76.0
77.8
78.9
71.7
72.9
64.9
65.0
70.1
70.1
58.9
58.1
65.0
69.6
66.0
70.4
PerceptionGPT-7B ( Pi et al., 2024 )
61.5
59.9
66.1
64.3
58.9
56.8
70.7
74.7
73.9
78.7
66.8
71.7
64.0
68.1
69.4
73.8
56.8
61.0
65.5
70.2
67.2
71.5
Table 1: Comparison with methods across GRES and RES benchmarks (%). Red, orange, gold, and yellow shading denote the first- through fourth-best scores in each column, respectively.
Figure 4: Qualitative comparison on challenging GRES examples. C o n t o u r V L A recovers multiple and part-level referents, preserves fine boundaries, and correctly rejects the no-target case.
Table 6
Figure 5: Contour evolution and latency on a subset of gRefCOCO val. (a) Boundary IoU (lines) and cumulative contour latency at steps 2, 4, and 8 (bars) for Full EASS and frozen observations. (b) Full EASS evolution from initial contours to final masks, with consistent instance colors.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Composition of the 139M LocateAnything-Data training instances used in Stage I. The mixture contains detection (66.9%), UI (16.5%), referring (7.3%), OCR (3.6%), layout (3.5%), and pointing (2.2%) data.
Stage
Training data
Supervision and objective
I. Grounding
LocateAnything-Data
Boxes and grounding responses; comprehensive grounding and detection pretraining.
II. Contour SFT
Four-dataset mixture
Boxes, masks, and aligned contour trajectories; joint grounding and geometric-action learning.
III. DECT-GRPO
Four-dataset mixture
Final segmentation feedback; joint optimization of the grounding and contour-action policies.
Appendix
Table 4: Complete training pipeline of C o n t o u r V L A . The four-dataset mixture comprises the training sets of gRefCOCO, RefCOCO, RefCOCO+, and RefCOCOg.
Figure 7: Qualitative comparisons on plural, no-target, and relational referring expressions.
Figure 8: Qualitative comparisons on numerical, spatial, and heterogeneous queries.
Figure 9: Qualitative segmentation on out-of-distribution videos. Panels (a)–(d) show sampled frames in temporal order, indicated by the arrows. The referring query is displayed at the left of each panel.
Component
Setting
Interpretation
Input / feature grid
896×896 / 56×56
Fixed-resolution input and stride-16 VLIM grid
Contour vertices
Pi∈[64,128]
Perimeter-adaptive valid-point count
Spatial offsets
J=4 , {oj}={1/64,1/32,1/16,1/8}
Fractions of box-normalized distance on each side
GAD depth
D=6 blocks
One routed semantic feature per block
Action horizon
K=2 chunks, T=4 steps/chunk
Eight action steps, with re-observation between the two chunks.
Action bound
δmax=0.05
Per-axis bound in normalized-box coordinates
Appendix
Table 5: Default spatial sampling, action execution, and optimization configuration. Stage-II and Stage-III batch sizes are global.
Routing variant
gIoU ↑
cIoU ↑
bIoU+↑
Predefined schedule
74.2
63.5
61.4
Slot-conditioned
74.7
64.2
62.2
Static-instance-conditioned
76.5
65.1
64.3
Full state-conditioned
77.6
67.4
66.0
Appendix
Table 6: State conditioning in semantic routing on gRefCOCO val (%). All variants retain SVS and spatial re-observation. gIoU and cIoU use the validation split; bIoU+ uses the fixed positive-target subset from Figure 5 at action step 8. The Full EASS boundary score is the Figure 5 endpoint ( 65.98% ), rounded to one decimal place. Bold denotes the best score in each column.
Table 7: Ablation of grounding and contour-action PG objectives (%). (a) End-to-end evaluation on the full gRefCOCO validation split using each model’s own grounding. (b) Eight-step contour evaluation on the fixed positive-target subset from Figure 5 , with the same SFT grounding and initial contours for every model. Bold denotes the best score in each column.
Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large language models (MLLMs) exhibit strong open-world visual grounding, but their outputs remain limited to sparse bounding-box coordinates and are insufficient for dense visual prediction. Recent MLLM-based segmentation methods either directly predict sparse contour coordinates, struggling to reconstruct continuous object boundaries, or rely on external segmentation foundation models such as the Segment Anything Model (SAM), introducing substantial architectural and deployment overhead. We present Qwen3-VL-Seg, a parameter-efficient framework that treats the MLLM-predicted box as a semantically grounded structural prior and decodes it into pixel-level referring segmentation. At its core, a lightweight box-guided mask decoder combines multi-scale spatial feature injection, spatial-semantic query construction, box-guided high-resolution pixel fusion, and iterative mask-aware query refinement, introducing only 17M parameters (about 0.4% of the base model). For scalable open-world training, we construct SA1B-ORS, an SA-1B-derived dataset with two subsets: SA1B-CoRS (category-oriented samples) and SA1B-DeRS (descriptive, instance-specific samples). For evaluation, we curate ORS-Bench, a manually screened benchmark with in-distribution and out-of-distribution subsets covering diverse referring expression types. Extensive experiments on referring expression segmentation, visual grounding, and ORS-Bench show that Qwen3-VL-Seg performs strongly across closed-set and open-world settings, with clear advantages on language-intensive instructions and strong out-of-distribution generalization. Evaluations on general multimodal benchmarks further show that the model broadly preserves general-purpose multimodal competence after segmentation-oriented adaptation.
While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visual shortcuts, associating actions with task-irrelevant visual features rather than the intended task semantics. These shortcuts block recomposition of elements already seen by the policy, that is, compositional generalization. Existing approaches mitigate such entanglement through task-relevant perception or targeted data diversification, but offer no explicit mechanism for unseen recomposition and require backbone-specific modifications with retraining. We observe that under such recomposition, VLAs often fail at global grounding while retaining local manipulation skills that recover near the correct target in familiar configurations. Therefore, we propose Referential Guidance (ReGuide), a training-free wrapper that, given object poses from a grounding module, combines semantic and geometric rebinding to guide the end-effector into demonstration-supported configurations of the instructed referent, where the frozen policy can resume execution. Experiments in simulation across multiple VLA backbones as well as on a real robot show that ReGuide improves success rates under compositional shifts by up to 56.8 and 75.0 percentage points, respectively, while preserving standard-task performance.
Yanyan Zhang, Disheng Liu, Xinpeng Li +8
Case Western Reserve University Cleveland, OH, USA
Referring expression segmentation requires language conditioned localization and pixel-accurate masks, but monolithic models can be costly to deploy. We present VespaSeg, a modular pipeline that grounds a text query with a compact vision-language model and converts the predicted box to a mask with MobileSAM. We study Florence-2-base, Florence-2-large, and Moondream2 grounders together with targeted adaptation of the grounding and segmentation stages. Under a repository-specific RefCOCO validation protocol containing the first expression for each of 3,811 referenced-object records, the adapted Florence-2-base pipeline obtains 73.64 mean intersection over union (mIoU) and 84.60 precision at IoU 0.5. On an NVIDIA RTX 6000 Ada GPU it processes 22.8 cached-image queries per second with 2.20 GB mean allocated GPU memory. A matched 500-query comparison gives 73.73 mIoU for Florence-2-base and 72.82 for Florence-2-large, while the base model is 1.70 times faster and uses 1.17 GB less allocated memory. Ablations show that ground-truth-box adaptation raises MobileSAM mIoU from 82.22 to 86.61 and that reducing the Florence-2 output-token budget from 64 to 32 preserves accuracy. These results support compact, modular grounding and segmentation, while also exposing the need for evaluation on the complete standard RefCOCO expression splits and deployment hardware.
Savindu Dilshan Wickramasinghe
University of Moratuwa · Department of Electronic and Telecommunication Engineering, University of Moratuwa Moratuwa, Sri Lanka