cs.CVSep 30, 2026

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Authors: Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, +14 more

Organizations: XPeng Inc. · Peking University · The University of Hong Kong · National University of Singapore · HKUST (GZ) · University of California, Berkeley · Princeton University

Abstract

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a 4.51×4.51\times speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.

Figures & tables

Explore similar work

CardsList
  1. LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

    May 26, 2026Shihao Wang, Shilong Liu, Yuanguo Kuang +10Vision-Language Model GroundingWeak Visual Grounding

  2. What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs

    May 10, 2026Jiaping Lin, Fei Shen, Junzhe Li +4Attention-Guided TrainingPrefill

  3. Grounded 3D-Aware Spatial Vision-Language Modeling

    May 28, 2026An-Chieh Cheng, Yang Fu, Yatai Ji +123D Visual GroundingSpatial Grounding