Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query--text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query--text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query--category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy--efficiency trade-off. The code will be released.
Figures & tables
Figure 1: Descriptions convey rich, reusable visual–semantic factors. Fine-grained cues learned from seen data may help match visual instances to unseen category concepts at test time.
Source
Samples
Regions
Coverage (%)
Cat.
Obj.
COCO
205,815
1,234,235
100.0
100.0
V3Det
166,106
683,814
100.0
100.0
GQA
358,181
3,972,578
93.2
97.1
Flickr30K Ent.
79,035
531,791
93.9
99.7
LLaVA-Cap
306,553
1,628,395
97.6
99.9
Table 1: Statistics of GroundingCapv2 after standardization.
Figure 2: Confidence-routed GroundingAgent-Opt. Qwen3-VL-8B agents curate GroundingCap-1M category and object annotations, routing uncertain cases to Qwen3-VL-32B. Together with inherited image descriptions, curated annotations form GroundingCapv2’s tri-level supervision.
Figure 3: Overview of RT-DETR-World. A frozen DINOv3 visual encoder, lightweight multi-scale projector, DETR decoder, and compact MiniLM form the deployable detector. DDA combines a deployment-consistent MiniLM pathway with training-only object- and image-level LLM supervision. RNR uses LLM-derived description similarities to relax semantically related negatives across the three description-alignment objectives. All teacher-side components are removed at inference.
Method
Backbone
Text Encoder
Training Data
FPS
AP/APFixed
APr/APrFixed
APc/APcFixed
APf/APfFixed
Parameters (M)
Visual/Det.
Text
Total
GLIP-T
Swin-T
BERT-base
OG
–
– / 24.9
– / 17.7
– / 19.5
– / 31.0
≈ 122
≈ 110
232
GLIP-T
Swin-T
BERT-base
OG, Cap4M
–
– / 26.0
– / 20.8
– / 21.4
– / 31.0
≈ 122
≈ 110
232
GLIPv2-T
Swin-T
BERT-base
OG, Cap4M
–
– / 29.0
– / –
– / –
– / –
≈ 122
≈ 110
232
Grounding DINO-T
Swin-T
BERT-base
OG
–
– / 25.6
– / 14.4
– / 19.6
– / 32.2
≈ 62
≈ 110
172
Grounding DINO-T
Swin-T
BERT-base
OG, Cap4M
–
– / 27.4
– / 18.1
– / 23.3
– / 32.7
≈ 62
≈ 110
172
Table 2: Zero-shot detection on LVIS minival, reported as standard AP / Fixed AP.
Method
Backbone
ODinW13
ODinW35
YOLO-Worldv2.1-S
YOLOv8-S
15.3
7.8
YOLO-Worldv2.1-M
YOLOv8-M
21.4
10.5
YOLO-Worldv2.1-L
YOLOv8-L
35.6
16.6
YOLOEv8-S
YOLOv8-S
28.0
12.7
YOLOEv8-M
YOLOv8-M
30.9
14.3
YOLOEv8-L
YOLOv8-L
31.9
14.8
Table 3: Zero-shot cross-domain transfer on ODinW.
Method
Backbone
COCO AP
COCO-O AP
ER
YOLO-Worldv2.1-S
YOLOv8-S
38.2
11.9
−5.3
YOLO-Worldv2.1-M
YOLOv8-M
43.8
16.7
−3.0
YOLO-Worldv2.1-L
YOLOv8-L
46.0
27.3
+6.6
YOLOEv8-S
YOLOv8-S
35.6
15.4
−0.6
YOLOEv8-M
YOLOv8-M
42.2
17.7
−1.3
YOLOEv8-L
YOLOv8-L
45.4
19.8
−0.6
Table 4: COCO-O distribution-shift robustness. ‡ indicates the use of COCO training images.
ID
Training Data
LLM Teacher
Deployable MiniLM Path
LVIS Minival
Image
Object
Category
Object Desc.
Source C/P
AP
AP r
AP c
AP f
B0
GroundingCap-1M
–
–
–
–
✓
27.9
18.3
24.8
32.4
B1
GroundingCap-1M
✓
✓
–
–
✓
32.9
23.6
30.2
36.9
C0
GroundingCapv2
✓
✓
✓
–
–
33.1
25.7
32.2
35.1
C1
GroundingCapv2
–
–
✓
✓
–
32.2
22.9
30.3
35.6
C2
GroundingCapv2
✓
–
✓
✓
–
33.1
27.8
32.6
34.4
Table 5: Ablation of training supervision on LVIS minival.
Auxiliary Supervision
AP
AP r
AP c
AP f
LLMDet-style Caption Generation
28.5
18.3
26.4
32.2
Hard Contrastive Alignment
33.6
27.6
31.7
36.4
SoftCLIP-style Soft Alignment
33.8
25.1
32.6
36.5
SRCL-style Similarity Reweighting
33.9
26.2
32.2
36.8
RNR (Ours)
35.2
32.0
33.8
36.9
Table 6: Auxiliary alignment strategies on LVIS minival.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Table 7: Hyperparameter sensitivity on LVIS minival. All results are reported using Fixed AP.
Figure 4: Qualitative zero-shot detections on LVIS minival. Using the full 1,203-category vocabulary, RT-DETR-World localizes diverse objects across indoor, outdoor, and cluttered scenes, including small and fine-grained instances. Labels and numbers indicate predicted categories and confidence scores, respectively.