Small object detection and tracking in videos remain critical yet underexplored challenges in computer vision, particularly for applications such as public safety, aerial surveillance, and autonomous driving. Existing benchmarks offer limited support due to limited numbers of small objects, constrained category diversity, and narrow scene coverage. To address these limitations, we introduce XS-VID, a large-scale video benchmark comprising 223K frames and 1.4M annotated bounding boxes across 374 video sequences spanning diverse scene types. XS-VID provides extensive coverage of small-object scales, particularly for extremely small (0∼122 pixels) and small (122∼202 pixels) objects, which collectively constitute over 55% of all annotations. For systematic evaluation, we establish three dedicated tracks: Detection, multiple object tracking (MOT), and single object tracking (SOT), and extensively test the existing state-of-the-art methods on each. The experimental results indicate that existing methods face significant challenges with XS-VID, mainly stemming from insufficient modeling of spatiotemporal features at small scales. To tackle these challenges, we propose a lightweight, high-precision detection framework dubbed YOLOFT. It enhances small-object feature representation and spatiotemporal integration while preserving high detection speed, thereby achieving improved accuracy and robustness on both the XS-VID and VisDrone benchmarks. Our dataset and code are publicly available at https://gjhhust.github.io/XS-VID/, providing a solid foundation for future research on small-object detection and tracking in videos.
Figures & tables
Fig. 1: Extremely small objects across diverse scenes in XS-VID. The dataset features challenging small-object detection scenarios with scale ratios as extreme as 1:40, spanning 16 scene types (11 illustrated here) including natural environments, urban areas, recreational zones, and infrastructure sites under day and night conditions. Bounding boxes highlight the minute scale of annotated objects, demonstrating XS-VID’s focus on extremely small-object detection challenges.
Fig. 2: AP-Latency comparison of various methods on our XS-VID. Our YOLOFT achieves the SOTA performance. Latency calculations are all with batch size = 1. Details in Tab. III .
Fig. 3: Showcases of our XS-VID dataset’s object size and challenges in small-object detection. (a) shows that the objects in our XS-VID dataset are extremely small, and (b) indicates that small-object detection mainly faces three challenges: high appearance variation (multi-appearance), semantic ambiguity between classes (mis-classification), and severe visual degradation caused by low resolution and insufficient texture details (texture distortion).
Dataset
0∼122 ( es )
122∼202 ( rs )
202∼322 ( gs )
Count
Ratio
Count
Ratio
Count
Ratio
ImageNetVID [ 6 ]
968
0.1%
13k
0.7%
47k
2.4%
UAVDT [ 1 ]
29k
3.3%
204k
23.0%
265k
29.8%
VisDrone VID [ 2 ]
7k
0.5%
115k
7.5%
316k
20.6%
XS-VID (ours)
418k
29.4%
370k
26.1%
237k
16.7%
TABLE I: Comparison of Small Object Area Distribution Across Various Datasets.
Dataset
Det.
Track.
Type
Domain
Seq.
Frames
Annotations
Resolution
FPS
T-Cat.
S-Cat.
Pub & Year
SIRST-v2 [ 19 ]
✓
IR
infrared
-
514
648
278 × 366
-
1
3
TGRS 2023
IRSTD-1K [ 20 ]
✓
IR
infrared
-
1K
1.5K
512 × 512
-
1
6
CVPR 2022
NUDT-SIRST [ 21 ]
✓
IR
infrared
-
1.3K
1.9K
256 × 256
-
1
5
TIP 2022
IRDST [ 27 ]
✓
IR
infrared
401
147K
144K
795 × 553
-
1
-
TGRS 2023
SIATD [ 28 ]
✓
IR
infrared
350
150K
247K
640 × 512
30
1
3
CSD 2021
Hui [ 29 ]
✓
✓
IR
infrared
22
16K
17K
256 × 256
100
2
-
CSD 2020
TABLE II: Comprehensive comparison of representative benchmarks for object detection, tracking, and joint detection and tracking (D&T) under small object scenarios. The Type column indicates the sensing modality, including RGB (visible), IR (infrared), and IR+RGB (dual-modality). T-Cat. denotes the number of target categories, and S-Cat. denotes the number of distinct scene types. Datasets are grouped by sensing modality (Type). Resolution denotes the released image size or the maximum image size for datasets with varying resolutions. XS-VID features 1.4 million annotations , which is comparable in scale to ImageNet VID (2 million), while focusing on extremely small objects. It combines extensive extremely small-object coverage with Detection, SOT, and MOT evaluation. Datasets marked with * are widely used but not specifically designed for small object detection. Varied means that the image sizes are not uniform.
Fig. 4: Quantitative Comparison between XS-VID and various datasets. (a) XS-VID exhibits a significantly higher proportion of extremely small ( 0∼122 ) and relatively small ( 122∼202 ) objects than other datasets. (b) Category distribution across XS-VID’s 7 classes. (c) Number of objects per frame in XS-VID is rich and balanced, compared to ImageNetVID and VisDrone. (d) The cumulative distribution of total object counts per video shows XS-VID emphasizes scene diversity with moderate video total target number.
Method
APtest
APestest
APrstest
APgstest
APmtest
APltest
#Params.
FLOPs
Latency [bs=1]
VOD
DFF [ 40 ]
10.9%
0.0%
0.2%
1.1%
7.6%
23.0%
62.12 M
187.24 G
20.9 ms
FGFA [ 41 ]
11.7%
0.0%
0.2%
1.5%
9.3%
23.4%
64.48 M
5579.10 G
151.5 ms
SELSA [ 42 ]
13.9%
0.0%
0.5%
2.3%
11.6%
28.1%
70.63 M
2846.46 G
88.5 ms
TROI [ 43 ]
12.3%
0.0%
0.3%
1.8%
10.4%
25.6%
78.12 M
3416.44 G
232.1 ms
MEGA [ 44 ]
9.8%
0.0%
0.1%
0.6%
6.8%
22.8%
–
–
–
DiffusionVID [ 45 ]
11.5%
0.0%
0.2%
1.4%
9.0%
23.3%
–
–
–
TABLE III: Comparison of representative image- and video-based detection methods on the XS-VID Detection Track. "#Params." refers to the number of model parameters. Method categories include VOD (video object detection), GOD (general object detection), SOD (small object detection), YOLO (YOLO series and variants), OVD (open-vocabulary detection), Unified (unified frameworks), and OURS (our method tailored for small object video detection, detailed in Section VII ). Temporal and image-based methods share the same frame-level evaluation protocol. We report the configurations listed in the Method column using their official implementations. ZS denotes zero-shot evaluation, and FT denotes fine-tuning on XS-VID. Reproduction details are provided in the SM’s Secs. II-A–II-B.
Fig. 5: Statistical analysis of single object tracking (SOT) trajectories in our dataset. (a) Histogram of track lengths and the corresponding average bounding box areas. The bars represent the number of trajectories within each length interval, while the red curve indicates the mean object size in each bin. (b) Spatial distribution heatmap of object centers aggregated across all frames. The relatively uniform coverage across the image space reflects diverse scene compositions and helps mitigate location bias.
Tracker
AUC (%)
Precision (%)
Speed (FPS)
Pub.
Year
Specialized trackers
SeqTrack b256 [ 103 ]
52.94
76.96
36
CVPR
2023
SeqTrack l384 [ 103 ]
54.26
78.34
7.6
CVPR
2023
PVT++ [ 107 ]
11.49
25.71
36.2
ICCV
2023
LoRAT b224 [ 108 ]
62.41
83.31
209
ECCV
2024
LoRAT g224 [ 108 ]
66.40
87.07
50
ECCV
2024
TABLE IV: Comparison on the XS-VID SOT track. AUC is the area under the success curve, and precision is measured at the 20-pixel threshold. Specialized trackers and unified frameworks are grouped separately.
Fig. 6: Trajectory-level statistics of the XS-VID MOT Track. (a) Scene-wise trajectory counts across 374 videos (T0–T307 training, E0–E65 testing) show variations in object density. (b) Occluded trajectories under different IoU thresholds. When IoU < 0.3 , RS-scale trajectories are more frequently occluded than GS or ES, increasing identity switch likelihood. (c) Trajectory duration vs. average bounding box area. Most trajectories cluster around 18 × 18 pixels , with long-duration trajectories (>250 frames) remaining in the small-object regime, highlighting the challenge of long-term association for tiny objects.
Fig. 7: Overview of the proposed YOLOFT architecture. The framework consists of three key components: (1) a lightweight 3D convolutional module for efficient local temporal modeling across adjacent video frames; (2) a Flow Mask Attention (FMA) module that dynamically predicts flow-guided offsets, attention weights, and fusion gates to improve temporal alignment and feature fusion; and (3) a mask-assisted auxiliary supervision branch that leverages motion-aware binary masks to provide targeted spatial supervision, enabling more efficient optimization and better localization of small objects.
Method
Paradigm
TAO mAP
MOTA
IDF1
(%)
(%)
(%)
Multi-stage with YOLOFT detections
BoostTrack [ 125 ]
T-by-D
14.4
17.3
40.3
SUSHI [ 127 ]
GNN
15.1
26.7
40.5
HybridSORT [ 134 ]
T-by-D
16.5
37.5
48.7
OC-SORT [ 122 ]
T-by-D
20.1
34.8
47.8
TABLE V: Expanded MOT evaluation on XS-VID. T-by-D, GNN, and E2E-Trans denote tracking by detection, graph neural network association, and end-to-end Transformer tracking, respectively. Detailed results and evaluation settings are provided in the SM’s Sec. II-D.
Fig. 8: The Flow Mask Attention (FMA) module combines temporal modeling with spatial feature refinement to adaptively enhance the current frame feature Ft . It uses a temporal encoder to capture short-term motion between Ft and the previous temporal feature F^t−1 , producing an updated temporal feature F^t .
Method
APval
APesval
APrsval
APgsval
APmval
APlval
#Params.
FLOPs
Latency [bs=1]
VOD
DFF [ 40 ]
10.3%
0.0%
0.1%
3.4%
13.6%
21.8%
120.0 M
458.7 G
25.5 ms
FGFA [ 41 ]
13.6%
0.0%
0.9%
6.3%
17.8%
28.5%
122.4 M
7874.6 G
181.8 ms
SELSA [ 42 ]
11.8%
0.0%
0.5%
2.7%
14.3%
30.2%
128.4 M
6917.7 G
110.0 ms
TROI [ 43 ]
12.0%
0.0%
0.1%
4.8%
16.6%
24.7%
136.0 M
7487.7 G
285.7 ms
TransVOD [ 47 ]
9.7%
1.0%
3.2%
4.9%
11.5%
23.8%
33.8 M
613.9 G
136.0 ms
StreamYOLO [ 46 ]
18.0%
1.6%
5.1%
10.6%
22.3%
33.9%
54.8 M
698.1 G
47.5 ms
TABLE VI: Performance comparison on VisDroneVID validation set. All AP values are reported in %. FLOPs measured at input size 6402 .
Method
AP test
AP estest
AP rstest
AP gstest
AP mtest
AP ltest
#Params.
FLOPs
Pub.
Year
Baseline
24.9%
9.5%
13.7%
14.7%
27.3%
40.2%
13.4 M
36.3 G
-
-
DWconv [ 139 ]
25.2%
9.5%
14.3%
13.4%
27.4%
43.1%
12.4 M
33.7 G
ICCV
2019
DYConv [ 140 ]
24.7%
10.2%
14.0%
14.0%
27.1%
39.6%
16.2 M
33.8 G
CVPR
2020
ODConv [ 141 ]
24.9%
9.4%
13.2%
13.7%
27.6%
41.5%
16.3 M
33.8 G
ICLR
2022
FDConv [ 137 ]
25.5%
9.5%
14.1%
14.6%
27.9%
41.6%
13.5 M
37.9 G
CVPR
2025
PConv [ 142 ]
25.3%
9.7%
14.0%
13.8%
27.5%
43.1%
13.4 M
36.2 G
AAAI
2025
TABLE VII: Exploration of convolution variants for small-object video detection. All AP values are reported in %. FLOPs are measured at input size 6402 .
Method
AP test
AP s
AP l
Latency [bs=1]
(%)
(%)
(%)
(ms)
Bilinear
23.9
12.4
39.1
20.5
CARAFE [ 143 ]
23.3
12.3
34.3
10.6
ConvTranspose
23.2
12.7
31.9
9.9
Deconv [ 144 ]
23.4
12.5
38.4
10.8
ESPCN [ 145 ]
24.5
13.7
37.1
11.2
TABLE VIII: Exploration of different upsampling methods in small object detection.
Loss Weight
AP test
AP s
AP m
AP l
(%)
(%)
(%)
(%)
baseline (GIoU [ 147 ] )
23.1
12.0
24.9
34.7
1.0*CIoU
23.2
12.4
24.5
33.6
1.0*NWD [ 148 ]
23.2
12.9
27.2
35.9
0.4CIoU + 0.6NWD
23.6
13.2
24.2
32.7
0.2CIoU + 0.8SAFit [ 4 ]
23.8
12.8
25.3
32.8
TABLE IX: Exploration of small-object localization losses on XS-VID. Mixed losses are weighted combinations of IoU-based and distribution-aware terms.
Method
AP test
AP stest
AP mtest
AP ltest
(%)
(%)
(%)
(%)
Baseline (YOLOv8)
22.7
12.6
23.7
36.0
+ Small object techniques
24.5
13.7
25.8
37.1
+ Flow Mask Attention (FMA)
24.3
13.0
27.1
37.6
+ Loss mask supervision
25.2
14.8
27.4
43.1
+ Video data augmentation
25.8
14.7
28.1
39.2
TABLE X: Ablation study on the final YOLOFT design. Each row incrementally adds one component to the YOLOv8 baseline to evaluate its contribution to small, medium, and large object detection. Bold values identify the final configuration rather than column-wise maxima.
[D, S, P]
AP test
AP es
AP rs
AP gs
(%)
(%)
(%)
(%)
[2.0, 0.2, 1e-4]
32.6
21.3
26.2
33.6
No augmentation
32.3
20.0
26.4
33.2
[0.0, 0.5, 1e-3]
32.1
20.8
25.3
33.1
[0.0, 0.0, 1e-3]
31.8
20.6
26.4
32.6
TABLE XI: Effect of different data augmentation strategies on the XS-VID snapshot (D: Degree, S: Scale, P: Perspective). Reproducibility details are provided in the SM’s Sec. II-F.
Method
AP test
AP s
FLOPs
Latency
(%)
(%)
(G)
(ms)
Concat
23.3
12.5
29.1
10.0
Add
23.6
12.3
28.8
9.0
ConvGRU
23.9
12.8
29.3
15.2
LSTM
23.5
12.4
31.7
18.4
3D Conv
24.7
13.4
23.7
13.9
TABLE XII: Ablation study on different temporal modeling strategies. FLOPs are measured at input size 6402 , and latency is tested with batch size 1 under FP16 precision.
History
Current
Weight
AP test (%)
–
–
–
23.2
–
✓
1.0
24.1
✓
✓
1.0
23.5
✓
✓
0.5
24.6
TABLE XIII: Ablation study on mask-assisted supervision. We vary whether historical and current frame masks are used, and the weight assigned to the auxiliary loss.
Fine-tuning on XS-VID under the corresponding official setting.
Added SOT trackers
XS-VID adaptation under the corresponding official task protocol.
MOT tracking-by-detection
Detector outputs come from the detector named in the table; association follows the corresponding tracker implementation.
Unified frameworks
Staged task adaptation following established unified-model practice [ 65 , 66 ] .
TABLE S2: Method-group summary of evaluation protocols for the added baselines.
Detector
Tracker / model
Paradigm
TAO Tracking mAP (%)
MOTChallenge
mAP
AP s
AP m
AP s∗
AP m∗
MOTA (%)
IDF1 (%)
IDSw
Multi-stage
YOLOFT
SUSHI [ 127 ]
GNN
15.1
1.4
15.1
9.7
27.5
26.7
40.5
4821
YOLOFT
BoostTrack [ 125 ]
T-by-D
14.4
8.0
11.6
9.9
14.6
17.3
40.3
9616
YOLOFT
HybridSORT [ 134 ]
T-by-D
16.5
4.0
16.2
11.9
26.6
37.5
48.7
2076
YOLOFT
OC-SORT [ 122 ]
T-by-D
20.1
8.8
26.1
11.6
32.0
34.8
47.8
7116
TABLE S3: Complete MOT evaluation on XS-VID. Multi-stage methods use YOLOFT detections and list the tracker and association paradigm separately; end-to-end and unified methods are reported as model-level entries. T-by-D, GNN, and E2E-Trans denote tracking by detection, graph neural network association, and end-to-end Transformer tracking, respectively.
Fig. S1: Qualitative comparison of YOLOFT with baseline detectors. YOLOFT produces predictions that align closely with ground-truth (GT) boxes across diverse scenarios. In case (a), YOLOFT captures the target size and position in a challenging small-object example. In case (c), it also detects medium-sized objects, indicating performance across object scales.
Pre-annotation input
Box AP (%)
AP es (%)
AP rs (%)
AP gs (%)
AP m (%)
AP l (%)
High resolution
10.9
3.0
9.5
6.6
30.5
35.2
Low resolution
2.2
0.0
0.1
0.5
9.4
10.1
TABLE S4: High- versus low-resolution pre-annotation using the official CO-DETR checkpoint.
Setting
AP (%)
AP es (%)
With ignore
29.3
11.5
Without ignore
28.6 ( −0.7 )
11.3 ( −0.2 )
TABLE S5: Effect of ignore-region handling on YOLOFT-L training. Values in parentheses denote changes relative to the setting with ignore-region handling.
Fig. S2: Representative XS-VID scenes illustrating variation in environment, viewpoint, illumination, target density, and extremely small-object appearance.
Small object detection remains a significant challenge due to feature degradation from downsampling, mutual occlusion in dense clusters, and complex background interference. To address these issues, this paper proposes FSDETR, a frequency-spatial feature enhancement framework built upon the RT-DETR baseline. By establishing a collaborative modeling mechanism, the method effectively leverages complementary structural information. Specifically, a Spatial Hierarchical Attention Block (SHAB) captures both local details and global dependencies to strengthen semantic representation. Furthermore, to mitigate occlusion in dense scenes, the Deformable Attention-based Intra-scale Feature Interaction (DA-AIFI) focuses on informative regions via dynamic sampling. Finally, the Frequency-Spatial Feature Pyramid Network (FSFPN) integrates frequency filtering with spatial edge extraction via the Cross-domain Frequency-Spatial Block (CFSB) to preserve fine-grained details. Experimental results show that with only 14.7M parameters, FSDETR achieves 13.9% APS on VisDrone 2019 and 48.95% AP50 tiny on TinyPerson, showing strong performance on small-object benchmarks. The code and models are available at https://github.com/YT3DVision/FSDETR.
Jianchao Huang, Fengming Zhang, Haibo Zhu +1
School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi, China
Small Object Detection (SOD) is a fundamental yet challenging problem in computer vision due to its limited spatial resolution and weak visual cues. Although recent approaches have achieved remarkable advances, the background distractors in different frequency spectra still degrade the performance. In this paper, we propose a novel small object detection framework termed SFDNet, which is capable of detecting small objects via efficient spectrum-aware feature disentanglement. Specifically, we propose an Adaptive Spectrum Disentanglement (ASD) module that decomposes backbone features into multiple complementary spectral components, aiming to construct discriminative object-relevant representations by discarding the background distractors for each component. Afterwards, to strengthen the semantic consistency of the similar objects in the same class, we propose a Class-Wise Prototype Distillation (CPD) procedure, which establishes class prototypes for the object instances and enforces the compact representation by efficient prototype distillation. Extensive experiments on multiple challenging benchmarks show that SFDNet outperforms existing state-of-the-art methods by a large margin. Our code is available at https://github.com/ManOfStory/SFDNet.
Yang Guo, Zihan Yang, Feifei Kou +3
Shenzhen Campus of Sun Yat-sen University · Beijing University of Posts and Telecommunications, Beijing, China · Hangzhou International Innovation Institute, Beihang University, Hangzhou, China +2
Small object detection (SOD) remains a challenging task in real-world applications. Despite recent advances, existing detectors remain limited by rigid processing that entangle spatial aggregation with implicit frequency aliasing and truncation, leading to inadequate preservation of high-frequency components for SOD. To tackle these limitations, we propose a Frequency-Spatial Domain Collaborative Detection Transformer (FSDC-DETR), a novel collaborative framework that explicitly models complementary spatial and frequency representations. Specifically, we first introduce Dual-Branch Frequency-Spatial Adaptive Fusion (DBFSAF) to enhance frequency diversity and adaptively capture frequency-spatial domain discriminative representations. Building on these representations, a frequency-spatial interaction scheme is further explored within the hybrid encoder to enable progressive feature propagation to the decoder. In particular, structure-aware frequency-spatial aggregation is achieved through Shunt Frequency-Spatial Feature Fusion (SFS-FF), establishing bidirectional interaction and progressive cross-scale propagation between frequency and spatial representations for coherent discriminative modeling. Meanwhile, informative high-frequency responses are preserved during scale transitions through Frequency-Spatial Dynamic Downsampling (FSD-Down), thereby minimizing frequency degradation throughout multi-scale fusion for the precise SOD. Experimental results demonstrate that FSDC-DETR achieves state-of-the-art performance, improving AP by 6.4 on VisDrone-DET2019 and 6.6 on AITODv2, with gains of 6.8 and 6.9 AP for small objects. The code is available at github.com/nevereverinsomnia/FSDC-DETR.
Aiwen Liu, Chengguang Zhu, Gang Wang +5
Micro-Intelligence, Shanghai 201100, China · East China Normal University, Shanghai 200241, China · Chongqing Normal University, Chongqing 401331, China +1