Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query--text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query--text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query--category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy--efficiency trade-off. The code will be released.
Figures & tables
Figure 1: Descriptions convey rich, reusable visual–semantic factors. Fine-grained cues learned from seen data may help match visual instances to unseen category concepts at test time.
Source
Samples
Regions
Coverage (%)
Cat.
Obj.
COCO
205,815
1,234,235
100.0
100.0
V3Det
166,106
683,814
100.0
100.0
GQA
358,181
3,972,578
93.2
97.1
Flickr30K Ent.
79,035
531,791
93.9
99.7
LLaVA-Cap
306,553
1,628,395
97.6
99.9
Table 1: Statistics of GroundingCapv2 after standardization.
Figure 2: Confidence-routed GroundingAgent-Opt. Qwen3-VL-8B agents curate GroundingCap-1M category and object annotations, routing uncertain cases to Qwen3-VL-32B. Together with inherited image descriptions, curated annotations form GroundingCapv2’s tri-level supervision.
Figure 3: Overview of RT-DETR-World. A frozen DINOv3 visual encoder, lightweight multi-scale projector, DETR decoder, and compact MiniLM form the deployable detector. DDA combines a deployment-consistent MiniLM pathway with training-only object- and image-level LLM supervision. RNR uses LLM-derived description similarities to relax semantically related negatives across the three description-alignment objectives. All teacher-side components are removed at inference.
Method
Backbone
Text Encoder
Training Data
FPS
AP/APFixed
APr/APrFixed
APc/APcFixed
APf/APfFixed
Parameters (M)
Visual/Det.
Text
Total
GLIP-T
Swin-T
BERT-base
OG
–
– / 24.9
– / 17.7
– / 19.5
– / 31.0
≈ 122
≈ 110
232
GLIP-T
Swin-T
BERT-base
OG, Cap4M
–
– / 26.0
– / 20.8
– / 21.4
– / 31.0
≈ 122
≈ 110
232
GLIPv2-T
Swin-T
BERT-base
OG, Cap4M
–
– / 29.0
– / –
– / –
– / –
≈ 122
≈ 110
232
Grounding DINO-T
Swin-T
BERT-base
OG
–
– / 25.6
– / 14.4
– / 19.6
– / 32.2
≈ 62
≈ 110
172
Grounding DINO-T
Swin-T
BERT-base
OG, Cap4M
–
– / 27.4
– / 18.1
– / 23.3
– / 32.7
≈ 62
≈ 110
172
Table 2: Zero-shot detection on LVIS minival, reported as standard AP / Fixed AP.
Method
Backbone
ODinW13
ODinW35
YOLO-Worldv2.1-S
YOLOv8-S
15.3
7.8
YOLO-Worldv2.1-M
YOLOv8-M
21.4
10.5
YOLO-Worldv2.1-L
YOLOv8-L
35.6
16.6
YOLOEv8-S
YOLOv8-S
28.0
12.7
YOLOEv8-M
YOLOv8-M
30.9
14.3
YOLOEv8-L
YOLOv8-L
31.9
14.8
Table 3: Zero-shot cross-domain transfer on ODinW.
Method
Backbone
COCO AP
COCO-O AP
ER
YOLO-Worldv2.1-S
YOLOv8-S
38.2
11.9
−5.3
YOLO-Worldv2.1-M
YOLOv8-M
43.8
16.7
−3.0
YOLO-Worldv2.1-L
YOLOv8-L
46.0
27.3
+6.6
YOLOEv8-S
YOLOv8-S
35.6
15.4
−0.6
YOLOEv8-M
YOLOv8-M
42.2
17.7
−1.3
YOLOEv8-L
YOLOv8-L
45.4
19.8
−0.6
Table 4: COCO-O distribution-shift robustness. ‡ indicates the use of COCO training images.
ID
Training Data
LLM Teacher
Deployable MiniLM Path
LVIS Minival
Image
Object
Category
Object Desc.
Source C/P
AP
AP r
AP c
AP f
B0
GroundingCap-1M
–
–
–
–
✓
27.9
18.3
24.8
32.4
B1
GroundingCap-1M
✓
✓
–
–
✓
32.9
23.6
30.2
36.9
C0
GroundingCapv2
✓
✓
✓
–
–
33.1
25.7
32.2
35.1
C1
GroundingCapv2
–
–
✓
✓
–
32.2
22.9
30.3
35.6
C2
GroundingCapv2
✓
–
✓
✓
–
33.1
27.8
32.6
34.4
Table 5: Ablation of training supervision on LVIS minival.
Auxiliary Supervision
AP
AP r
AP c
AP f
LLMDet-style Caption Generation
28.5
18.3
26.4
32.2
Hard Contrastive Alignment
33.6
27.6
31.7
36.4
SoftCLIP-style Soft Alignment
33.8
25.1
32.6
36.5
SRCL-style Similarity Reweighting
33.9
26.2
32.2
36.8
RNR (Ours)
35.2
32.0
33.8
36.9
Table 6: Auxiliary alignment strategies on LVIS minival.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Table 7: Hyperparameter sensitivity on LVIS minival. All results are reported using Fixed AP.
Figure 4: Qualitative zero-shot detections on LVIS minival. Using the full 1,203-category vocabulary, RT-DETR-World localizes diverse objects across indoor, outdoor, and cluttered scenes, including small and fine-grained instances. Labels and numbers indicate predicted categories and confidence scores, respectively.
Open-domain open-vocabulary detection (ODOVD) requires detectors to generalize to both novel categories and unseen domains, making it more challenging than open-vocabulary detection. Existing methods typically train open-vocabulary detectors together with domain generalization modules from scratch, leading to high training cost. we propose ExDet, a lightweight category-domain collaborative generalization framework for ODOVD that enhances the cross-category and cross-domain generalization of existing detectors. ExDet consists of Text-Guided Extrapolation (TGE), a lightweight Detector-Compatible Rectification (DCR) module, and ExRPN. Specifically, TGE exploits the DeltaSpace property of vision-language models (VLMs) to infer category- and domain-aware proxy visual prototypes from text. DCR is learned from the TGE-generated prototypes in a detector training-free and real-data-free manner, and is inserted after the classification head at inference to rectify representations toward a detector-compatible source-domain visual distribution, thereby enhancing classification for targets from novel categories and unseen domains. ExRPN recalibrates proposal scores by combining semantic similarity with RPN confidence, improving recall for novel and domain-shifted objects while providing better support for subsequent classification and DCR. ExDet achieves SOTA performance on OD-LVIS, OV-LVIS, Objects365, and MSOSB.
Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World frequently misalign vague language with the intended objects. Building on our prior work Commonsense-Guided Open-World Object Detection Using LLMs and Visual-Semantic Matching, we address YOLO-World's limitations in grounding task-driven queries. We propose Vague2Detect, a hybrid pipeline in which a fine-tuned Sentence-BERT retrieves candidates from a structured household Knowledge Base (KB), and YOLO-World verifies their presence in the image. For prompts outside the KB, a large language model (GPT-3.5-turbo) generates candidate descriptions, dynamically expanding the KB to cover novel concepts. On a benchmark of household scenes using custom images and an Open Images V7 subset, YOLO-World alone achieves only 32% Vague Prompt Success Rate (VPSR), the ability to map ambiguous queries to correct detections. In contrast, Vague2Detect improves performance to 61% VPSR with high precision, and up to 85% when augmented with GPT fallback.
Ibrohimjon Muminov, Jihie Kim
Dongguk University, Seoul, South Korea · Department of Computer Science and AI, Dongguk University, Seoul, South Korea
Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-world category spaces continually evolve, and existing OVD models still struggle with newly emerging concepts, while repeated full retraining is prohibitively expensive. To this end, we introduce a new task setting, termed Continual OVD with Novel Concept Injection (COVD), where models sequentially learn incoming novel concept groups while preserving prior concepts and original open-vocabulary knowledge, along with a new benchmark, Novel-114. Our key observation is that pretrained visual encoders often already perceive and represent many novel concepts, and the main bottleneck lies in the lack of stable semantic alignment between visual representations and textual concepts. Based on this, we propose NoIn-Det, an efficient continual injection framework without additional parameters. NoIn-Det freezes the visual encoder, preserves the text representation space using only texts of common concepts and previously injected concepts, and injects novel concepts by updating only a small subset of text-branch parameters beneficial to novel concept learning. Extensive experiments show that NoIn-Det effectively learns novel concepts, preserves old knowledge, and consistently outperforms existing continual learning methods for VLMs without introducing additional parameters.Novel-114 and the code will be released.
Yupeng Zhang, Ruize Han, Yuzhong Feng +3
Tianjin University. · Shenzhen University of Advanced Technology.