GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Organizations: XPeng Inc. · Peking University · The University of Hong Kong · Tsinghua University · National University of Singapore · HKUST (GZ) · University of California, Berkeley
Abstract
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.
Figures & tables
| 3 6 8 15 16 19 Model | Common | Long-tailed | Dense & Tiny | Referring grounding | ||||
|---|---|---|---|---|---|---|---|---|
| COCO | LVIS | Dense200 | VisDrone | HumanRef | RefCOCOg val | RefCOCOg test | RefCOCO avg | |
| Closed-set Specialized Detectors | ||||||||
| DINO-R50 * ( Zhang et al., 2022 ) | 55.60 | – | – | – | – | – | – | – |
| DETR-R50 * ( Carion et al., 2020 ) | 48.30 | – | – | – | – | – | – | – |
| Open-set Specialized Detectors | ||||||||
| GroundingDINO ( Liu et al., 2024 ) | 60.56 | 52.61 | 24.92 | 34.47 | 46.13 | 49.77 | 50.43 | 45.15 |
| 3 5 15 16 22 Model | Robot and spatial pointing | GUI grounding | ||||
|---|---|---|---|---|---|---|
| RefSpatial (avg) | RefSpatial Unseen | RoboSpatial Context | ScreenSpot-Pro | ScreenSpot-V2 | OSWorld-G | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 14.25 | 4.33 | 4.92 | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||
| JEDI * ( Xie et al., 2026 ) | – | – | – | 36.10 | 88.60 | – |
| UI-R1 * ( Lu et al., 2026 ) | – | – | – | 17.80 | 85.40 | – |
| 3 6 13 14 17 Model | OCR | Layout grounding | ||||
|---|---|---|---|---|---|---|
| HierText | ICDAR2015 | TotalText | SROIE | DocLayNet | M6Doc | |
| Closed-set Specialized Detectors | ||||||
| DocLayout-YOLO * ( Zhao et al., 2024 ) | – | – | – | – | 81.10 | – |
| PaddleOCRv5 * ( Cui et al., 2025 ) | 30.50 | 25.60 | 25.70 | 58.60 | – | – |
| Vision-Language Models (<10B) | ||||||
| Qwen3-VL-4B ( Bai et al., 2025a ) | 23.48 | 28.41 | 38.35 | 40.41 | 40.81 | 24.73 |
| 3 5 13 14 17 Model | Referring object pointing | Common / long-tailed | Dense / tiny | ||||
|---|---|---|---|---|---|---|---|
| HumanRef | RefCOCOg val | RefCOCOg test | COCO | LVIS | Dense200 | VisDrone | |
| Open-set Specialized Detectors | |||||||
| GroundingDINO ( Liu et al., 2024 ) | 46.85 | 49.34 | 49.97 | 70.41 | 55.07 | 32.93 | 39.45 |
| Vision-Language Models (<10B) | |||||||
| Qwen3-VL-4B ( Bai et al., 2025a ) | 66.89 | 76.43 | 77.64 | 65.33 | 55.08 | 21.72 | 23.50 |
| Qwen3.5-9B ( Qwen Team, 2026a ) | 78.21 | 77.59 | 77.85 | 72.21 | 64.00 | 65.35 | 44.73 |
| 4 | COCO | Dense200 | ||||
|---|---|---|---|---|---|---|
| Model | Boxes/img | Tokens/img | Tokens/box | Boxes/img | Tokens/img | Tokens/box |
| SEED1.5-VL ( Guo et al., 2025 ) | 4.2 | 631.0 | 148.8 | 73.1 | 5446.3 | 74.5 |
| GroundingPI | 6.0 | 45.6 | 7.6 | 87.2 | 444.7 | 5.1 |
| Token or delimiter | Role |
|---|---|
| <0> – <999> | One atomic token per quantized coordinate |
| </c> | Separator between categories in the input query |
| <|object_ref_start|> , <|object_ref_end|> | Delimit a semantic label or transcription |
| <|box_start|> , <|box_end|> | Delimit a box, point, or ordered-point payload |
| None | Ordinary text indicating an absent queried target |
| <|im_end|> | End of the assistant response |
| Task | User prompt | Output |
|---|---|---|
| Grounding / dense grounding | Locate all the instances that match the following categories: cat1 </c> cat2 </c> …. | Boxes |
| Referring | Locate the target referred to by the following description: phrase . | Boxes |
| Object pointing | Point to: cat1 </c> cat2 </c> …. | Points |
| Referring pointing | Point to the target referred to by the following description: phrase . | Points |
| GUI grounding | Point to the UI element to click for the following instruction: instruction . | Point |
| OCR | OCR task detect all the text in box format. | Text + boxes |
| Component | Configuration |
|---|---|
| Vision encoder | 27 layers; hidden width 1024; FFN width 4096; 12 attention heads; QKV hidden width 1536; patch size 14 |
| Spatial aggregation | neighboring patches; four 1024-dimensional features form one 4096-dimensional feature |
| Projector | LayerNorm(1024), spatial aggregation, two bias-free linear layers ( ) with GELU, and RMSNorm(2560); normalization |
| Language backbone | 36 full-attention layers; hidden width 2560; FFN width 9728; 32 query heads and 8 KV heads; head dimension 128 |
| Language numerics | SiLU; attention dropout 0; RMSNorm ; RoPE base ; no sliding window |
| Positional capacity | 262,144 positions |
| Base VLM training (Pretrain 1) | Pretrain 2 / SFT | |||
|---|---|---|---|---|
| Setting | Stage 1: Projector alignment | Stage 2: Joint multimodal pretraining | Stage 3: General visual/video understanding | Stage 4: Coordinate alignment |
| Trainable modules | Projector only | All | All | All |
| Per-device batch | 6 | 2 | 2 | 2 |
| Global batch | 768 | 256 | 256 | 256 |
| Context / packing length | 8192 / 8192 | 8192 / 8192 | 32768 / 32768 | 8192 / 8192 |
| Language peak LR | Frozen | |||
| Setting | Value |
|---|---|
| GPUs / precision / sharding | 256 / BF16 / ZeRO-2 |
| Responses per prompt | 8 |
| Per-device response batch / accumulation | 1 / 1 |
| Distinct-prompt / response batch | 32 / 256 |
| Updates per rollout | 1 |
| Trainable modules | Language parameters, including embedding and output head |
| Component | Weight |
| Structured-format validity | 0.10 |
| Agreement between predicted and reference counts | 0.10 |
| Mean detection F1 at IoU , , and | 0.50 |
| Matched-box IoU | 0.25 |
| Compliance with the output ordering convention | 0.05 |
| Oversized-box penalty |
| Annotation structure | Global views | Content score |
|---|---|---|
| With granularity conflict | ||
| Without granularity conflict |
| Component | Shared configuration |
|---|---|
| Backbone conditions | 8 normalized-depth hidden states |
| Condition interface | 64 tokens per depth, width 1024 |
| Action topology | 16 atomic blocks: 8 cross-attention and 8 self-attention blocks, interleaved from cross-attention |
| Residual / attention width | 1024; 16 heads; head dimension 64 |
| Feed-forward width | 4096 |
| Planning tokens | 32 |
| Item | Configuration |
|---|---|
| Optimization length | 40,000 updates |
| Hardware / precision | 40 accelerators; bfloat16; DeepSpeed ZeRO-2 ( Rajbhandari et al., 2020 ) |
| Batch size | 32 per device; 1280 global |
| Optimizer | AdamW ( Loshchilov & Hutter, 2019 ) , , , |
| Weight decay / clipping | ; gradient norm 1.0 |
| Learning rates | for the action expert and interface; for adapted backbone parameters |
| 3 7 9 26 27 35 Model | Zero-shot | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | ||
| Closed-set Specialized Detectors | ||||||||||
| DINO-R50 * ( Zhang et al., 2022 ) | NO | 62.60 | 76.50 | 68.80 | 17.80 | 25.80 | 21.10 | 50.00 | 62.40 | 55.60 |
| DETR-R50 * ( Carion et al., 2020 ) | NO | 59.60 | 73.90 | 65.90 | 10.60 | 19.00 | 13.60 | 42.90 | 55.30 | 48.30 |
| DyHead-R50 * ( Dai et al., 2021 ) | NO | 58.10 | 76.60 | 66.10 | 11.90 | 20.60 | 15.00 | 44.80 | 60.10 | 51.30 |
| Open-set Specialized Detectors | ||||||||||
| 3 5 22 23 31 Model | Zero-shot | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | ||
| Open-set Specialized Detectors | ||||||||||
| GroundingDINO ( Liu et al., 2024 ) | YES | 52.64 | 82.61 | 64.30 | 23.34 | 26.55 | 24.84 | 44.30 | 59.58 | 52.61 |
| Vision-Language Models (<10B) | ||||||||||
| Rex-Omni ( Jiang et al., 2026 ) | YES | 58.50 | 73.11 | 64.99 | 17.76 | 20.04 | 18.83 | 44.02 | 46.58 | 46.74 |
| LocateAnything Fast ( Wang et al., 2026a ) | YES | 48.06 | 61.58 | 53.98 | 20.37 | 25.12 | 22.50 | 38.43 | 48.55 | 42.90 |
| 3 5 22 23 31 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 22.43 | 36.30 | 27.73 | 10.92 | 18.38 | 13.70 | 20.16 | 32.62 | 24.92 |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 70.61 | 75.13 | 72.80 | 8.66 | 9.17 | 8.91 | 51.79 | 54.88 | 53.29 |
| LocateAnything Fast ( Wang et al., 2026a ) | 22.45 | 26.12 | 24.15 | 9.16 | 10.50 | 9.78 | 19.30 | 22.31 | 20.70 |
| 3 5 22 23 31 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 32.38 | 87.09 | 47.21 | 2.81 | 7.42 | 4.08 | 23.65 | 63.54 | 34.47 |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 42.47 | 52.33 | 46.89 | 1.07 | 1.27 | 1.16 | 24.78 | 30.12 | 27.19 |
| LocateAnything Fast ( Wang et al., 2026a ) | 13.58 | 15.44 | 14.45 | 0.85 | 0.97 | 0.90 | 9.27 | 10.44 | 9.82 |
| 3 5 22 23 31 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 54.80 | 46.81 | 50.49 | 37.38 | 31.26 | 34.05 | 50.32 | 42.59 | 46.13 |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 85.91 | 84.95 | 85.43 | 65.72 | 65.09 | 65.40 | 80.33 | 79.42 | 79.87 |
| LocateAnything Fast ( Wang et al., 2026a ) | 68.01 | 71.34 | 69.64 | 53.25 | 55.15 | 54.19 | 62.15 | 64.72 | 63.41 |
| 3 5 22 23 31 Model | RefCOCOg val | RefCOCOg test | ||||
|---|---|---|---|---|---|---|
| F1@.50 | F1@.95 | F1mIoU | F1@.50 | F1@.95 | F1mIoU | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 58.37 | 22.47 | 49.77 | 58.52 | 24.38 | 50.43 |
| Vision-Language Models (<10B) | ||||||
| Rex-Omni ( Jiang et al., 2026 ) | 87.01 | 35.23 | 73.90 | 87.36 | 36.51 | 74.76 |
| LocateAnything Fast ( Wang et al., 2026a ) | 87.99 | 39.39 | 75.30 | 88.34 | 42.08 | 76.50 |
| 3 5 22 23 30 Model | RefCOCO | RefCOCOg | RefCOCOplus | ||||||
|---|---|---|---|---|---|---|---|---|---|
| F1@.50 | F1@.95 | F1mIoU | F1@.50 | F1@.95 | F1mIoU | F1@.50 | F1@.95 | F1mIoU | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 51.19 | 23.28 | 44.33 | 57.99 | 23.17 | 49.69 | 49.20 | 20.89 | 41.42 |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 84.37 | 32.85 | 71.43 | 84.74 | 34.74 | 72.34 | 76.92 | 29.53 | 64.10 |
| LocateAnything Fast ( Wang et al., 2026a ) | 92.48 | 43.95 | 80.61 | 88.47 | 40.69 | 76.17 | 84.43 | 39.53 | 73.08 |
| 3 5 23 24 32 Model | HumanRef | RefCOCOg val | RefCOCOg test | ||
|---|---|---|---|---|---|
| R@Point | P@Point | F1@Point | F1@Point | F1@Point | |
| Open-set Specialized Detectors | |||||
| GroundingDINO ( Liu et al., 2024 ) | 50.22 | 43.90 | 46.85 | 49.34 | 49.97 |
| Vision-Language Models (<10B) | |||||
| Rex-Omni ( Jiang et al., 2026 ) | 84.05 | 82.76 | 83.40 | 84.96 | 85.32 |
| LocateAnything Fast ( Wang et al., 2026a ) | 66.10 | 68.28 | 67.18 | 73.84 | 74.94 |
| 3 5 23 24 32 Model | COCO | LVIS | ||||
|---|---|---|---|---|---|---|
| R@Point | P@Point | F1@Point | R@Point | P@Point | F1@Point | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 68.92 | 71.97 | 70.41 | 44.91 | 71.15 | 55.07 |
| Vision-Language Models (<10B) | ||||||
| Rex-Omni ( Jiang et al., 2026 ) | 77.81 | 81.77 | 79.74 | 63.46 | 78.14 | 70.04 |
| LocateAnything Fast ( Wang et al., 2026a ) | 68.99 | 74.26 | 71.53 | 55.81 | 70.05 | 62.13 |
| 3 5 23 24 32 Model | Dense200 | VisDrone | ||||
|---|---|---|---|---|---|---|
| R@Point | P@Point | F1@Point | R@Point | P@Point | F1@Point | |
| Open-set Specialized Detectors | ||||||
| GroundingDINO ( Liu et al., 2024 ) | 22.00 | 65.45 | 32.93 | 27.20 | 71.76 | 39.45 |
| Vision-Language Models (<10B) | ||||||
| Rex-Omni ( Jiang et al., 2026 ) | 75.59 | 77.76 | 76.66 | 47.81 | 56.92 | 51.97 |
| LocateAnything Fast ( Wang et al., 2026a ) | 64.63 | 66.65 | 65.63 | 55.22 | 57.55 | 56.36 |
| 3 5 24 25 36 Model | Point-in-mask accuracy | |||
|---|---|---|---|---|
| RefSpatial Location | RefSpatial Placement | RefSpatial Unseen | RoboSpatial Context | |
| Open-set Specialized Detectors | ||||
| GroundingDINO ( Liu et al., 2024 ) | 26.50 | 2.00 | 4.33 | 4.92 |
| Vision-Language Models (<10B) | ||||
| Rex-Omni ( Jiang et al., 2026 ) | 51.00 | 52.50 | 37.01 | 59.02 |
| LocateAnything Fast ( Wang et al., 2026a ) | 53.00 | 19.00 | 16.88 | 15.57 |
| 3 5 7 24 25 33 Model | HierText | ICDAR2015 | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| F1@.50 | F1@.75 | F1@.95 | F1mIoU | Parse err. | F1@.50 | F1@.75 | F1@.95 | F1mIoU | Parse err. | |
| Closed-set Specialized Detectors | ||||||||||
| PaddleOCRv5 * ( Cui et al., 2025 ) | 45.20 | – | 3.40 | 30.50 | – | 38.20 | – | 1.20 | 25.60 | – |
| Open-set Specialized Detectors | ||||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||||||
| 3 5 7 24 25 33 Model | TotalText | SROIE | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| F1@.50 | F1@.75 | F1@.95 | F1mIoU | Parse err. | F1@.50 | F1@.75 | F1@.95 | F1mIoU | Parse err. | |
| Closed-set Specialized Detectors | ||||||||||
| PaddleOCRv5 * ( Cui et al., 2025 ) | 40.20 | – | 0.70 | 25.70 | – | 77.70 | – | 5.60 | 58.60 | – |
| Open-set Specialized Detectors | ||||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||||||
| 3 5 25 26 34 Model | Dev | Creative | CAD | Sci | Office | OS | Overall | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Text | Icon | Text | Icon | Text | Icon | Text | Icon | Text | Icon | Text | Icon | Action acc. | Parse err. | |
| Open-set Specialized Detectors | ||||||||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 61.04 | 9.66 | 53.03 | 12.59 | 23.35 | 9.38 | 57.64 | 26.36 | 65.54 | 24.53 | 42.06 | 13.48 | 36.75 | – |
| LocateAnything Fast ( Wang et al., 2026a ) | 68.18 | 44.83 | 59.60 | 32.87 | 58.38 | 35.94 | 71.53 | 53.64 | 75.14 | 56.60 | 51.40 | 37.08 | 56.04 | – |
| 3 5 25 26 33 Model | ScreenSpot-V2 | OSWorld-G | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mobile text | Mobile icon | Desktop text | Desktop icon | Web text | Web icon | Action acc. | Parse err. | Exact acc. | Parse err. | |
| Open-set Specialized Detectors | ||||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | ||||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 96.90 | 82.46 | 97.94 | 80.71 | 89.74 | 76.35 | 88.29 | – | 46.10 | – |
| LocateAnything Fast ( Wang et al., 2026a ) | 94.48 | 81.04 | 93.81 | 86.43 | 88.89 | 83.25 | 88.44 | – | 59.93 | – |
| 3 5 7 24 25 33 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Closed-set Specialized Detectors | |||||||||
| DocLayout-YOLO * ( Zhao et al., 2024 ) | – | – | 91.20 | – | – | 52.10 | – | – | 81.10 |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | |||||||||
| 3 5 22 23 31 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 73.09 | 78.27 | 75.59 | 18.16 | 18.78 | 18.46 | 53.36 | 56.64 | 54.95 |
| LocateAnything Fast ( Wang et al., 2026a ) | 59.82 | 61.54 | 60.67 | 18.06 | 18.50 | 18.28 | 46.39 | 47.69 | 47.03 |
| 3 5 22 23 30 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 7.94 * | 20.16 * | 11.39 * | 1.56 * | 4.94 * | 2.37 * | 6.34 * | 16.01 * | 9.08 * |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 78.97 | 77.92 | 78.44 | 9.13 | 9.01 | 9.07 | 57.56 | 56.75 | 57.15 |
| LocateAnything Fast ( Wang et al., 2026a ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| 3 5 22 23 30 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 2.03 * | 9.98 * | 3.38 * | 1.35 * | 5.64 * | 2.18 * | 1.90 * | 9.00 * | 3.14 * |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 73.05 | 74.81 | 73.92 | 10.89 | 11.12 | 11.00 | 54.93 | 56.08 | 55.50 |
| LocateAnything Fast ( Wang et al., 2026a ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| 3 5 22 23 30 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 26.84 * | 27.30 * | 27.07 * | 13.69 * | 13.52 * | 13.60 * | 23.62 * | 23.83 * | 23.73 * |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 78.69 | 67.17 | 72.48 | 20.38 | 18.49 | 19.39 | 61.99 | 53.44 | 57.40 |
| LocateAnything Fast ( Wang et al., 2026a ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |
| 3 5 22 23 30 Model | IoU 0.50 | IoU 0.95 | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R | P | F1 | R | P | F1 | R | P | F1 | |
| Open-set Specialized Detectors | |||||||||
| GroundingDINO ( Liu et al., 2024 ) | 18.16 * | 18.90 * | 18.52 * | 9.85 * | 9.99 * | 9.92 * | 15.91 * | 16.38 * | 16.14 * |
| Vision-Language Models (<10B) | |||||||||
| Rex-Omni ( Jiang et al., 2026 ) | 69.31 | 61.02 | 64.90 | 17.42 | 15.42 | 16.36 | 52.74 | 46.39 | 49.36 |
| LocateAnything Fast ( Wang et al., 2026a ) | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * | N/A * |