RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
Authors: Zijun Liao, Yian Zhao, Xin Shan, Yu Yan, Chang Liu, Lei Lu, Xiangyang Ji, Jie Chen
Organizations: School of Electronic and Computer Engineering, Peking University, Shenzhen, China · Department of Automation and BNRist, Tsinghua University, Beijing, China
Real-time object detection has achieved substantial progress through meticulously designed architectures and optimization strategies. However, the pursuit of high-speed inference via lightweight network designs often leads to degraded feature representation, which hinders further performance improvements and practical on-device deployment. In this paper, we propose a cost-effective and highly adaptable distillation framework that harnesses the rapidly evolving capabilities of Vision Foundation Models (VFMs) to enhance lightweight object detectors. Given the significant architectural and learning objective disparities between VFMs and resource-constrained detectors, achieving stable and task-aligned semantic transfer is challenging. To address this, on one hand, we introduce a \textbf{Deep Semantic Injector (DSI)} module that facilitates the integration of high-level representations from VFMs into the deep layers of the detector. On the other hand, we devise a \textbf{Gradient-guided Adaptive Modulation (GAM)} strategy, which dynamically adjusts the intensity of semantic transfer based on gradient norm ratios. Without increasing deployment and inference overhead, our approach painlessly delivers striking and consistent performance gains across diverse DETR-based models, underscoring its practical utility for real-time detection. Our new model family, RT-DETRv4, achieves state-of-the-art results on COCO, attaining AP scores of 49.8/53.7/55.4/57.0 at corresponding speeds of 273/169/124/78 FPS. Code is publicly available at https://github.com/RT-DETRs/RT-DETRv4.
Figures & tables
(a) COCO AP val vs. T4 GPU Latency.
(b) COCO AP val vs. Params (M).
Figure 1 : Overview of RT-DETRv4 . We leverage a Vision Foundation Model (VFM) to extract high-quality semantic representations, which are aligned with the deepest feature map ( F5 ) from the AIFI module via a Feature Projector in the Deep Semantic Injector (DSI). To ensure faster and more stable convergence, a Gradient-guided Adaptive Modulation (GAM) dynamically adjusts the DSI loss during training. The proposed framework operates only during the training phase (dashed arrows and blue blocks) of the real-time detector and keeps the original architecture unchanged during inference and deployment, introducing no additional overhead while improving accuracy.
Figure 2 : Semantic injection for object detection. (a) Illustration from the information bottleneck perspective. (b) Visualization of nuisance details in images versus robust semantic representations from a VFM.
Figure 3 : Illustration of different Deep Semantic Injector (DSI) strategies. (a) Direct alignment of multi-scale backbone features ( S3,S4,S5 ). (b) Hybrid alignment of both backbone and the AIFI features ( S3,S4,S5,F5 ). (c) Ours: Targeted alignment of only the AIFI feature ( F5 ), which possesses the highest-level semantics. This design allows gradients to backpropagate, enhancing both the AIFI module and the backbone.
Figure 4 : Comparison of dense features. We compare the feature map quality of DEIM-L (top) and RT-DETRv4-L (bottom) by projecting dense outputs to RGB space using PCA. The visualization reveals that our DSI module substantially enhances the semantic representation of deep features. From left to right: input image, AIFI feature map F5 , and CCFF features P3 , P4 , P5 .
Method
AP val
AP 50val
AP 75val
RT-DETRv2-L
52.1
70.2
56.7
w/ DSI
52.3 (+0.2)
70.4 (+0.2)
56.4 (-0.3)
w/ DSI+GAM
52.6 (+0.5)
70.7 (+0.5)
56.8 (+0.1)
D-FINE-L
53.1
70.8
57.4
w/ DSI
53.2 (+0.1)
70.8 (+0)
57.7 (+0.3)
w/ DSI+GAM
53.4 (+0.3)
71.1 (+0.3)
58.0 (+0.6)
Table 2 : Results of ablation on DSI and GAM across multiple detectors. Our method brings consistent and significant gains with zero additional inference cost.
Position
S3
S4
S5
F5
AP val
AP 50val
AP 75val
Baseline
-
-
-
-
53.8
71.4
58.5
Backbone
✓
-
-
-
53.7
71.2
58.4
-
✓
-
-
53.7
71.3
58.4
-
-
✓
-
53.8
71.3
58.5
✓
✓
✓
-
53.7
71.3
58.5
Hybrid
✓
✓
✓
✓
53.8
71.4
58.4
Table 3 : Results of ablation on semantic injection position. We compare the effectiveness of applying DSI at different positions of DEIM-L , corresponding to the strategies in Figure 3 . Aligning only the AIFI output ( F5 ) yields the best performance.
Projector Arch.
AP val
AP 50val
AP 75val
DEIM-L
53.8
71.4
58.5
w/ 1x1 Conv
53.8
71.5
58.5
w/ MLP
54.2
71.7
59.0
w/ Linear
54.3 (+0.5)
71.8 (+0.4)
59.0 (+0.5)
Table 4: Results of ablation on the feature projector.
Semantic Teacher
AP val
AP 50val
AP 75val
DEIM-L
53.8
71.4
58.5
w/ MAE-ViT-B
53.9
71.4
58.6
w/ DINOv2-ViT-Base
54.1
71.6
58.8
w/ DINOv3-ViT-Base
54.3 (+0.5)
71.9 (+0.5)
59.0 (+0.5)
Table 5 : Results of ablation on the semantic teacher.
Loss Function
AP val
AP 50val
AP 75val
DEIM-M
52.5
69.9
57.2
w/ MSE Loss
52.7
70.0
57.4
w/ Cosine Similarity
53.7 (+1.2)
71.0 (+1.1)
58.4 (+1.2)
Table 6 : Results of ablation on the alignment loss. The cosine similarity loss demonstrates superior performance. DEIM-M is adopted as the baseline, with all models trained for 90 epochs and 12 EMA epochs.
λ
0.1
0.2
0.4
1
2
4
10
20
30
50
100
GAM
AP val
54.7
54.7
54.8
54.9
54.9
55.0
55.0
55.1
55.0
54.9
54.6
55.4
Table 7 : Ablation on the loss weighting strategy. GAM consistently outperforms the static weighting. DEIM-L is adopted as the baseline, with all models trained for 50 epochs and 8 EMA epochs.
Figure 5 : AP val evolution on COCO during training.
Figure A : Visualization of RT-DETRv4-L predictions. Comparison with novel real-time detectors including YOLOv12 [ 34 ] , YOLOv13 [ 18 ] , D-FINE [ 27 ] , and DEIM [ 16 ] . Red circles indicate missed detections, while white circles denote duplicate detections or misclassifications.
Figure B : Visualization of the deepest feature F5 using PCA. We compare the semantic representations of the deepest feature F5 between the baseline DEIM [ 16 ] (without semantic injection) and our method (trained with semantic injection but without retaining semantic teacher at inference).
Module
Epoch
Baseline
DSI ( λ=2 )
DSI ( λ=20 )
DSI+GAM
Avg Abs Norm
%
Avg Abs Norm
%
Avg Abs Norm
%
Avg Abs Norm
%
AIFI
1-12
1.5033
0.64%
2.7968
1.22%
6.0068
2.55%
3.9483
1.71%
12-24
1.1364
0.53%
2.5365
1.17%
4.6953
2.13%
4.2419
1.94%
24-36
1.0300
0.48%
2.2927
1.08%
3.8318
1.81%
4.0486
1.89%
36-48
0.9829
0.45%
2.2152
1.03%
3.5430
1.67%
4.1631
1.94%
48-58
1.0628
0.45%
2.5098
1.08%
4.2570
1.85%
3.3276
1.46%
Table A : Detailed hyper-parameters of RT-DETRv4. The training hyper-parameters for RT-DETRv4 models, focusing on the settings that diverge from the baselines (RT-DETRv2 [ 24 ] , D-FINE [ 27 ] , and DEIM [ 16 ] ).
Figure 17
ρ%,δ%
1, 0
1.5, 0
1.8, 0
2, 0
2.2, 0
2.5, 0
1, 0.1
1.5, 0.1
1.8, 0.1
2, 0.1
2.2, 0.1
2.5, 0.1
AP val
55.0
55.1
55.2
55.2
55.1
55.0
55.0
55.1
55.2
55.4
55.2
55.1
ρ%,δ%
1, 0.25
1.5, 0.25
1.8, 0.25
2, 0.25
2.2, 0.25
2.5, 0.25
1, 0.5
1.5, 0.5
1.8, 0.5
2, 0.5
2.2, 0.5
2.5, 0.5
AP val
54.9
55.1
55.1
55.3
55.2
55.2
54.9
55.0
55.0
55.2
55.1
55.0
Table D : Ablation on GAM hyper-parameters. DEIM-L is adopted as the baseline, with all models trained for 50 epochs and 8 EMA epochs.
Feature map
S3
S4
S5
F5
CKA similarity
0.05
0.30
0.40
0.59
Table E : CKA similarity between detector layers and VFM (DINOv3).