GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
Organizations: XPeng Inc. · Peking University · The University of Hong Kong · National University of Singapore · HKUST (GZ) · University of California, Berkeley · Princeton University
Abstract
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.
Figures & tables
| 3 6 8 16 17 18 21 Model | Common | Long-tailed | Dense & Tiny | Referring grounding | |||
|---|---|---|---|---|---|---|---|
| COCO | LVIS | Dense200 | VisDrone | RefCOCOg val | RefCOCOg test | RefCOCO avg | |
| Closed-set Specialized Detectors | |||||||
| DINO-R50 * ( Zhang et al., 2022 ) | 55.60 | – | – | – | – | – | – |
| DETR-R50 * ( Carion et al., 2020 ) | 48.30 | – | – | – | – | – | – |
| Open-set Specialized Detectors | |||||||
| GroundingDINO ( Liu et al., 2024 ) | 60.56 | 52.61 | 24.92 | 34.47 | 49.77 | 50.43 | 45.15 |
| 3 5 16 17 18 24 Model | Robot and spatial pointing | GUI grounding | ||||
|---|---|---|---|---|---|---|
| RefSpatial (avg) | RefSpatial Unseen | RoboSpatial Context | ScreenSpot-Pro | ScreenSpot-V2 | OSWorld-G | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 14.25 | 4.33 | 4.92 | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||
| JEDI * ( Xie et al., 2026 ) | – | – | – | 36.10 | 88.60 | – |
| UI-R1 * ( Lu et al., 2026 ) | – | – | – | 17.80 | 85.40 | – |
| 3 6 14 15 16 19 Model | OCR | Layout grounding | Visual prompting | |||||
|---|---|---|---|---|---|---|---|---|
| HierText | ICDAR2015 | TotalText | SROIE | DocLayNet | M6Doc | FSC147 | Dense200 | |
| Closed-set Specialized Detectors | ||||||||
| DocLayout-YOLO * ( Zhao et al., 2024 ) | – | – | – | – | 81.10 | – | – | – |
| PaddleOCRv5 * ( Cui et al., 2025 ) | 30.50 | 25.60 | 25.70 | 58.60 | – | – | – | – |
| Vision-Language Models (<10B) | ||||||||
| Qwen3-VL-4B ( Bai et al., 2025a ) | 23.48 | 28.41 | 38.35 | 40.41 | 40.81 | 24.73 | 18.90 | 2.49 |
| 3 5 14 15 16 19 Model | Referring object pointing | Common / long-tailed | Dense / tiny | |||
|---|---|---|---|---|---|---|
| RefCOCOg val | RefCOCOg test | COCO | LVIS | Dense200 | VisDrone | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 49.34 | 49.97 | 70.41 | 55.07 | 32.93 | 39.45 |
| Vision-Language Models (<10B) | ||||||
| Qwen3-VL-4B ( Bai et al., 2025a ) | 76.43 | 77.64 | 65.33 | 55.08 | 21.72 | 23.50 |
| Qwen3.5-9B ( Qwen Team, 2026a ) | 77.59 | 77.85 | 72.21 | 64.00 | 65.35 | 44.73 |
| Task | Instruction |
|---|---|
| Category grounding | Locate all the instances that match the following categories: [CATEGORIES]. |
| Referring grounding | Locate the target referred to by the following description: [EXPRESSION]. |
| Category pointing | Point to: [CATEGORIES]. |
| Referring pointing | Point to the target referred to by the following description: [EXPRESSION]. |
| GUI grounding | Point to the UI element to click for the following instruction: [INSTRUCTION]. |
| OCR | OCR task detect all the text in box format. |
| Token(s) | Function |
|---|---|
| <|object_ref_start|> , <|object_ref_end|> | Delimit a semantic identifier |
| <|box_start|> , <|box_end|> | Delimit a geometric payload |
| <0> – <999> | Atomic quantized coordinates |
| </c> | Separate requested categories |
| <|vision_start|> , <|vision_end|> | Delimit visual input |
| <|image_pad|> | Image placeholder in the input template |
| Component | Configuration |
|---|---|
| Vision encoder | 27 layers; width 1024; FFN width 4096; 12 heads; QKV width 1536; patch size 14 |
| Projector | LayerNorm(1024), aggregation, bias-free linear layers with GELU, then RMSNorm(2560); normalization |
| Language backbone | 36 layers; width 2560; FFN width 9728; 32 query heads and 8 KV heads; head dimension 128 |
| Language numerics | SiLU; attention dropout 0; RMSNorm ; RoPE base |
| Positional capacity | 262,144 positions |
| Training phase | GroundAnything-VLM | GroundAnything |
|---|---|---|
| Base VLM Training (Pretrain 1) | Stages 1–3; causal cross-entropy | Stages 1–3; causal cross-entropy |
| Coordinate Alignment (Pretrain 2 / SFT) | Stage 4; assistant-only spatial cross-entropy | Stage 4; assistant-only spatial cross-entropy |
| AR-to-Diffusion conversion | — | ; shared weights |
| Reinforcement learning (RL) | Causal GRPO | Causal GRPO; shared AR/diffusion weights |
| Base VLM Training (Pretrain 1) | Pretrain 2 | |||
|---|---|---|---|---|
| Setting | Stage 1 | Stage 2 | Stage 3 | Stage 4 |
| Trainable modules | Projector | All | All | All |
| Language peak LR | Frozen | |||
| Vision peak LR | Frozen | |||
| Projector peak LR | ||||
| Per-device batch / accumulation | 6 / 1 | 2 / 1 | 2 / 1 | 2 / 1 |
| Setting | Phase I | Phase II |
|---|---|---|
| GPUs | 256 | 256 |
| Per-device / global batch | 32 / 8192 | 32 / 8192 |
| Clean sequence length | 4096 | 4096 |
| Trainable modules | Projector and language | All |
| Language / projector peak LR | / | / |
| Vision peak LR | Frozen |
| Setting | GroundAnything | GroundAnything-VLM |
|---|---|---|
| GPUs | 256 | 256 |
| Per-device response batch / accumulation | 1 / 1 | 1 / 1 |
| Responses per prompt | 8 | 8 |
| Prompt / response batch | 32 / 256 | 32 / 256 |
| Updates per rollout | 1 | 1 |
| Learning rate |
| Branch | Parse status | Grade |
| Word/line | Complete, valid structure with no invalid instances | 1.00 |
| Incomplete structure with reliably parsed instances | 0.60 | |
| Only some valid instances can be recovered | 0.30 | |
| No reliable parse | 0.00 | |
| Complementary | Strictly valid structure with no invalid content | 1.00 |
| references | Complete main structure with extraneous content | 0.45 |
| 3 7 9 26 27 28 36 Model | Zero-shot | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | ||
| Closed-set Specialized Detectors | ||||||||||
| DINO-R50 * ( Zhang et al., 2022 ) | NO | 62.60 | 76.50 | 68.80 | 17.80 | 25.80 | 21.10 | 50.00 | 62.40 | 55.60 |
| DETR-R50 * ( Carion et al., 2020 ) | NO | 59.60 | 73.90 | 65.90 | 10.60 | 19.00 | 13.60 | 42.90 | 55.30 | 48.30 |
| DyHead-R50 * ( Dai et al., 2021 ) | NO | 58.10 | 76.60 | 66.10 | 11.90 | 20.60 | 15.00 | 44.80 | 60.10 | 51.30 |
| Open-set Specialized Detectors | ||||||||||
| 3 5 22 23 24 32 Model | Zero-shot | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | ||
| Open-set Specialized Detectors | ||||||||||
| GroundingDINO ( Liu et al., 2024 ) | YES | 52.64 | 82.61 | 64.30 | 23.34 | 26.55 | 24.84 | 44.30 | 59.58 | 52.61 |
| Vision-Language Models (<10B) | ||||||||||
| Rex-Omni ( Jiang et al., 2026 ) | YES | 58.50 | 73.11 | 64.99 | 17.76 | 20.04 | 18.83 | 44.02 | 46.58 | 46.74 |
| LocateAnything Fast ( Wang et al., 2026 ) | YES | 48.06 | 61.58 | 53.98 | 20.37 | 25.12 | 22.50 | 38.43 | 48.55 | 42.90 |
| 3 5 22 23 24 32 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 22.43 | 36.30 | 27.73 | 10.92 | 18.38 | 13.70 | 20.16 | 32.62 | 24.92 |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 70.61 | 75.13 | 72.80 | 8.66 | 9.17 | 8.91 | 51.79 | 54.88 | 53.29 |
| LocateAnything Fast ( Wang et al., 2026 ) | 22.45 | 26.12 | 24.15 | 9.16 | 10.50 | 9.78 | 19.30 | 22.31 | 20.70 |
| 3 5 22 23 24 32 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 32.38 | 87.09 | 47.21 | 2.81 | 7.42 | 4.08 | 23.65 | 63.54 | 34.47 |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 42.47 | 52.33 | 46.89 | 1.07 | 1.27 | 1.16 | 24.78 | 30.12 | 27.19 |
| LocateAnything Fast ( Wang et al., 2026 ) | 13.58 | 15.44 | 14.45 | 0.85 | 0.97 | 0.90 | 9.27 | 10.44 | 9.82 |
| 3 5 22 23 24 32 Model | RefCOCOg val | RefCOCOg test | ||||
|---|---|---|---|---|---|---|
| F1@.50 | F1@.95 | F1mIoU | F1@.50 | F1@.95 | F1mIoU | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 58.37 | 22.47 | 49.77 | 58.52 | 24.38 | 50.43 |
| Vision-Language Models (<10B) | ||||||
| Rex-Omni ( Jiang et al., 2026 ) | 87.01 | 35.23 | 73.90 | 87.36 | 36.51 | 74.76 |
| LocateAnything Fast ( Wang et al., 2026 ) | 87.99 | 39.39 | 75.30 | 88.34 | 42.08 | 76.50 |
| 3 5 22 23 24 31 Model | RefCOCO | RefCOCOg | RefCOCOplus | ||||||
|---|---|---|---|---|---|---|---|---|---|
| F1@.50 | F1@.95 | F1mIoU | F1@.50 | F1@.95 | F1mIoU | F1@.50 | F1@.95 | F1mIoU | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 51.19 | 23.28 | 44.33 | 57.99 | 23.17 | 49.69 | 49.20 | 20.89 | 41.42 |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 84.37 | 32.85 | 71.43 | 84.74 | 34.74 | 72.34 | 76.92 | 29.53 | 64.10 |
| LocateAnything Fast ( Wang et al., 2026 ) | 92.48 | 43.95 | 80.61 | 88.47 | 40.69 | 76.17 | 84.43 | 39.53 | 73.08 |
| 3 5 23 24 25 33 Model | RefCOCOg val | RefCOCOg test |
|---|---|---|
| F1@Point | F1@Point | |
| Open-set Specialized Detectors | ||
| GroundingDINO ( Liu et al., 2024 ) | 49.34 | 49.97 |
| Vision-Language Models (<10B) | ||
| Rex-Omni ( Jiang et al., 2026 ) | 84.96 | 85.32 |
| LocateAnything Fast ( Wang et al., 2026 ) | 73.84 | 74.94 |
| 3 5 23 24 25 33 Model | COCO | LVIS | ||||
|---|---|---|---|---|---|---|
| R@Point | P@Point | F1@Point | R@Point | P@Point | F1@Point | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 68.92 | 71.97 | 70.41 | 44.91 | 71.15 | 55.07 |
| Vision-Language Models (<10B) | ||||||
| Rex-Omni ( Jiang et al., 2026 ) | 77.81 | 81.77 | 79.74 | 63.46 | 78.14 | 70.04 |
| LocateAnything Fast ( Wang et al., 2026 ) | 68.99 | 74.26 | 71.53 | 55.81 | 70.05 | 62.13 |
| 3 5 23 24 25 33 Model | Dense200 | VisDrone | ||||
|---|---|---|---|---|---|---|
| R@Point | P@Point | F1@Point | R@Point | P@Point | F1@Point | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 22.00 | 65.45 | 32.93 | 27.20 | 71.76 | 39.45 |
| Vision-Language Models (<10B) | ||||||
| Rex-Omni ( Jiang et al., 2026 ) | 75.59 | 77.76 | 76.66 | 47.81 | 56.92 | 51.97 |
| LocateAnything Fast ( Wang et al., 2026 ) | 64.63 | 66.65 | 65.63 | 55.22 | 57.55 | 56.36 |
| 3 5 24 25 26 37 Model | Point-in-mask accuracy | |||
|---|---|---|---|---|
| RefSpatial Location | RefSpatial Placement | RefSpatial Unseen | RoboSpatial Context | |
| Open-set Specialized Detectors | ||||
| GroundingDINO ( Liu et al., 2024 ) | 26.50 | 2.00 | 4.33 | 4.92 |
| Vision-Language Models (<10B) | ||||
| Rex-Omni ( Jiang et al., 2026 ) | 51.00 | 52.50 | 37.01 | 59.02 |
| LocateAnything Fast ( Wang et al., 2026 ) | 53.00 | 19.00 | 16.88 | 15.57 |
| 3 5 7 24 25 26 34 Model | HierText | ICDAR2015 | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| F1@.50 | F1@.75 | F1@.95 | F1mIoU | Parse err. | F1@.50 | F1@.75 | F1@.95 | F1mIoU | Parse err. | |
| Closed-set Specialized Detectors | ||||||||||
| PaddleOCRv5 * ( Cui et al., 2025 ) | 45.20 | – | 3.40 | 30.50 | – | 38.20 | – | 1.20 | 25.60 | – |
| Open-set Specialized Detectors | ||||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||||||
| 3 5 7 24 25 26 34 Model | TotalText | SROIE | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| F1@.50 | F1@.75 | F1@.95 | F1mIoU | Parse err. | F1@.50 | F1@.75 | F1@.95 | F1mIoU | Parse err. | |
| Closed-set Specialized Detectors | ||||||||||
| PaddleOCRv5 * ( Cui et al., 2025 ) | 40.20 | – | 0.70 | 25.70 | – | 77.70 | – | 5.60 | 58.60 | – |
| Open-set Specialized Detectors | ||||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||||||
| 3 5 25 26 27 35 Model | Dev | Creative | CAD | Sci | Office | OS | Overall | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Text | Icon | Text | Icon | Text | Icon | Text | Icon | Text | Icon | Text | Icon | Action acc. | Parse err. | |
| Open-set Specialized Detectors | ||||||||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 61.04 | 9.66 | 53.03 | 12.59 | 23.35 | 9.38 | 57.64 | 26.36 | 65.54 | 24.53 | 42.06 | 13.48 | 36.75 | – |
| LocateAnything Fast ( Wang et al., 2026 ) | 68.18 | 44.83 | 59.60 | 32.87 | 58.38 | 35.94 | 71.53 | 53.64 | 75.14 | 56.60 | 51.40 | 37.08 | 56.04 | – |
| 3 5 25 26 27 34 Model | ScreenSpot-V2 | OSWorld-G | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mobile text | Mobile icon | Desktop text | Desktop icon | Web text | Web icon | Action acc. | Parse err. | Exact acc. | Parse err. | |
| Open-set Specialized Detectors | ||||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 96.90 | 82.46 | 97.94 | 80.71 | 89.74 | 76.35 | 88.29 | – | 46.10 | – |
| LocateAnything Fast ( Wang et al., 2026 ) | 94.48 | 81.04 | 93.81 | 86.43 | 88.89 | 83.25 | 88.44 | – | 59.93 | – |
| 3 5 7 24 25 26 34 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Closed-set Specialized Detectors | |||||||||
| DocLayout-YOLO * ( Zhao et al., 2024 ) | – | – | 91.20 | – | – | 52.10 | – | – | 81.10 |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | |||||||||
| 3 5 22 23 24 32 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 73.09 | 78.27 | 75.59 | 18.16 | 18.78 | 18.46 | 53.36 | 56.64 | 54.95 |
| LocateAnything Fast ( Wang et al., 2026 ) | 59.82 | 61.54 | 60.67 | 18.06 | 18.50 | 18.28 | 46.39 | 47.69 | 47.03 |
| 3 5 22 23 24 31 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 7.94 * | 20.16 * | 11.39 * | 1.56 * | 4.94 * | 2.37 * | 6.34 * | 16.01 * | 9.08 * |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 78.97 | 77.92 | 78.44 | 9.13 | 9.01 | 9.07 | 57.56 | 56.75 | 57.15 |
| LocateAnything Fast ( Wang et al., 2026 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| 3 5 22 23 24 31 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 2.03 * | 9.98 * | 3.38 * | 1.35 * | 5.64 * | 2.18 * | 1.90 * | 9.00 * | 3.14 * |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 73.05 | 74.81 | 73.92 | 10.89 | 11.12 | 11.00 | 54.93 | 56.08 | 55.50 |
| LocateAnything Fast ( Wang et al., 2026 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |