OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing Images
Authors: Zeyu Ding, Yong Zhou, Jiaqi Zhao, Hancheng Zhu, Wen-Liang Du, Rui Yao, Abdulmotaleb El Saddik
Organizations: School of Computer Science and Technology, China University of Mining and Technology, and with Mine Digitization Engineering Research Center of the Ministry of Education, and also with Innovation Research Center of Disaster Intelligent Prevention and Emergency Rescue, China University of Mining and Technology, Xuzhou 221116, China · School of Electrical Engineering and Computer Science, University of Ottawa, Ottawa, ON K1N 6N5, Canada
Oriented object detection in remote sensing images is a challenging task due to objects being distributed in multi-orientation. Recently, end-to-end transformer-based methods have achieved success by eliminating the need for post-processing operators compared to traditional CNN-based methods. However, directly extending transformers to oriented object detection presents three main issues: 1) objects rotate arbitrarily, necessitating the encoding of angles along with position and size; 2) the geometric relations of oriented objects are lacking in self-attention, due to the absence of interaction between content and positional queries; and 3) oriented objects cause misalignment, mainly between values and positional queries in cross-attention, making accurate classification and localization difficult. In this paper, we propose an end-to-end transformer-based oriented object detector, consisting of three dedicated modules to address these issues. First, Gaussian positional encoding is proposed to encode the angle, position, and size of oriented boxes using Gaussian distributions. Second, Wasserstein self-attention is proposed to introduce geometric relations and facilitate interaction between content and positional queries by utilizing Gaussian Wasserstein distance scores. Third, oriented cross-attention is proposed to align values and positional queries by rotating sampling points around the positional query according to their angles. Experiments on six datasets DIOR-R, a series of DOTA, HRSC2016 and ICDAR2015 show the effectiveness of our approach. Compared with previous end-to-end detectors, the OrientedFormer gains 1.16 and 1.21 AP50 on DIOR-R and DOTA-v1.0 respectively, while reducing training epochs from 3× to 1×. The codes are available at https://github.com/wokaikaixinxin/ai4rs and https://github.com/wokaikaixinxin/OrientedFormer.
Figures & tables
Fig. 1: (a) Object instances distribute in remote sensing images with arbitrary orientation. Angles are used to characterize oriented objects, in addition to positionsa and sizes. (b) Visualization of sampling points of the oriented cross-attention for alignment.
Notation
Description
Notation
Description
(x,y,w,h,θ)
oriented boxes
dq
dimension of self-attention
(x,y,z,r,θ)
oriented boxes
G
Wasserstein distance score
{fl}l=1L
levels of feature map
τ , ϵ
coefficient of G
Hl
height of feature map
(Δx,Δy,Δz)
offset of sampling points
Wl
width of feature map
g
number of heads
D
channel of feature map
O
number of sampling points
TABLE I: Nomenclature with related notations.
Fig. 2: Overall architecture of the OrientedFormer. Features are extracted from images. An object query is decomposed into a content query Qc and a positional query Qp . The Gaussian PE encodes positional queries. The Wasserstein self-attention measures the geometric relations between two different content queries by utilizing Wasserstein distance scores. The oriented cross-attention is proposed to align values and positional queries.
Fig. 3: An example of Gaussian positional encoding. (a) positional encoding in Deformable DETR. (b) Gaussian positional encoding.
Fig. 4: Self-attention in the decoder. (a) vanilla self-attention. (b)Wasserstein self-attention.
Fig. 5: Oriented cross-attention. It attends to sparse sampling points (x~,y~,z~) around the center of a positional query. Sampling points are rotated according to angles for alignment. Values V are interpolated by sampling points and multi-scale features. We deploy attention mechanisms separately on each particular dimension of values, i.e., scale-aware, channel-aware, and spatial-aware.
Method
config
value
OrientedFormer
optimizer
AdamW
base learning rate
5e-5
weight decay
1e-6
optimizer momentum
β1,β2 =0.9, 0.999
batch size
4
GPUs
2
TABLE II: Experiment settings of OrientedFormer.
Mehtod
Backbone
APL
APO
BF
BC
BR
CH
DAM
ETS
ESA
GF
GTF
HA
OP
SH
STA
STO
TC
TS
VE
WM
AP 50
one-stage:
RetinaNet-O [ 33 ]
R50
61.49
28.52
73.57
81.17
23.98
72.54
19.94
72.39
58.20
69.25
79.54
32.14
44.87
77.71
67.57
61.09
81.46
47.33
38.01
60.24
57.55
DFDet [ 40 ]
R50
61.92
38.83
77.41
81.36
34.11
74.97
26.26
62.31
76.06
75.56
79.62
38.26
52.76
80.40
73.11
68.27
81.38
52.23
44.11
63.35
62.11
Oriented Rep [ 9 ]
R50
70.03
46.11
76.12
87.19
39.14
78.76
34.57
71.80
80.42
76.16
79.41
45.48
54.90
87.82
77.03
68.07
81.60
56.83
51.57
71.25
66.71
DCFL [ 7 ]
R50
68.60
53.10
76.70
87.10
42.10
78.60
34.50
71.50
80.80
79.70
79.50
47.30
57.40
85.20
64.60
66.40
81.50
58.90
50.90
70.90
66.80
two-stage:
TABLE III: Comparison with state-of-the-art methods on the DIOR-R . The results in bold denote the best performance of each column.
Method
Backbone
PL
BD
BR
GTF
SV
LV
SH
TC
BC
ST
SBF
RA
HA
SP
HC
AP 50
one-stage:
PSC [ 17 ]
R50
88.24
74.42
48.63
63.44
79.98
80.76
87.59
90.88
82.02
71.58
59.12
60.78
65.78
71.21
53.06
71.83
R3Det [ 15 ]
R101
88.76
83.09
50.91
67.27
76.23
80.39
86.72
90.78
84.68
83.24
61.98
61.35
66.91
70.63
53.94
73.79
S 2 A-Net [ 16 ]
R50
89.11
82.84
48.37
71.11
78.11
78.39
87.25
90.83
84.90
85.64
60.36
62.60
65.26
69.13
57.94
74.12
H2RBox [ 44 ]
R50
88.93
78.89
46.27
68.79
81.12
75.45
86.68
90.89
86.71
87.33
64.15
68.83
62.81
69.39
59.79
74.40
CFA [ 18 ]
R50
88.34
83.09
51.92
72.23
79.95
78.68
87.25
90.90
85.38
85.71
59.63
63.05
73.33
70.36
47.86
74.51
TABLE IV: Comparison with state-of-the-art methods on the DOTA-v1.0 dataset. * denotes multi-scale trainig and testing. The results in bold denote the best performance of each column.
Fig. 6: P-R curve (IoU=0.5) on the DIOR-R test set with ResNet50.
Method
SV
LV
SH
ST
SP
CC
AP 50
RetinaNet-O [ 33 ]
44.53
56.79
73.31
59.96
64.52
0.83
59.16
Faster RCNN-O [ 53 ]
51.28
68.98
79.37
67.50
65.28
1.54
62.00
Mask R-CNN [ 54 ]
51.31
71.34
79.75
66.07
64.46
9.42
62.67
HTC [ 55 ]
51.54
73.31
80.31
67.34
64.48
5.15
63.40
ReDet [ 5 ]
52.38
75.73
80.92
68.64
70.55
11.53
66.86
OrientedFormer
64.05
77.04
85.33
78.11
72.08
10.86
67.06
TABLE V: Main results of small size objects on DOTA-v1.5 . The results in bold denote the best performance of each column.
Method
Precision
Recall
F-measure
FLOPs
SASM [ 8 ]
56.7
77.9
65.7
71G
PSC [ 17 ]
83.7
63.2
72.0
78G
Retinanet-O [ 33 ]
83.9
67.5
74.8
77G
GWD [ 56 ]
84.4
67.6
75.1
77G
R3DET [ 15 ]
83.2
69.2
75.6
120G
Oriented RCNN [ 4 ]
73.9
80.9
77.2
100G
TABLE VI: Main results of precision, recall (IoU=0.5), F-measure and Flops on ICDAR2015 . The results in bold denote the best performance. The backbone used by all methods is Resnet50.
Fig. 7: P-R curve (IoU=0.75) on the DIOR-R test set with ResNet50.
Methods
Backbone
mAP(07)
mAP(12)
PSC [ 17 ]
R50
85.65
-
RoI Transformer [ 41 ]
R101
86.20
-
Gliding Vertex [ 6 ]
R101
88.20
-
PIoU [ 57 ]
DLA34
89.20
-
CenterMap [ 58 ]
R50
-
92.8
R3Det [ 15 ]
R101
89.26
96.01
TABLE VII: Comparison with the state-of-the-art methods on the HRSC2016 dataset. 07 means evaluation under pascal voc2007 metric, and 12 means evaluation under pascal voc2012 metric.
Method
SASM [ 8 ]
RetinaNet-O [ 33 ]
Oriented Rep [ 9 ]
Mask R-CNN [ 54 ]
AP 50
44.53
46.68
48.95
49.47
Method
ATSS-O [ 61 ]
S 2 A-Net [ 16 ]
HTC [ 55 ]
DCFL [ 7 ]
AP 50
49.57
49.86
50.34
51.57
Method
RoI Trans. [ 41 ]
S 2 A-Net + DCFL
Oriented R-CNN [ 4 ]
OrientedFormer
AP 50
52.81
52.84
53.28
54.27
TABLE VIII: Performance comparisons on the DOTA-v2.0 dataset.
TABLE IX: OrientedFormer ablation experiments with ResNet-50 on DIOR-R. Default choice for our model is colored gray
PE
-
Deform.
DAB.
Learnable
Gaussian
AP 50
66.85
65.97
66.03
64.27
67.28
TABLE X: Comparsions of different positional encoding in the decoder.
Methods
Gaussian PE
Wasserstein Self-Attention
Oriented Cross-Attention
DIOR-R
AP 50
AP 75
AP 50:95
Oriented Former
62.69
44.18
41.38
✓
✓
63.08
43.44
41.00
✓
65.78
43.69
41.87
✓
✓
66.85
46.22
43.73
✓
✓
67.03
44.07
42.49
TABLE XI: The effectiveness of proposed individual modules on DIOR-R.
Self-Attention
-
iof
iou
Wasserstein
AP 50
67.03
66.57
67.08
67.28
TABLE XII: Comparsions of different self-attention in the decoder.
Methods
Gaussian PE
Wasserstein Self-Attention
Oriented Cross-Attention
DOTA-v1.0
AP 50
AP 75
AP 50:95
Oriented Former
73.81
47.74
45.40
✓
✓
74.55
49.26
46.28
✓
74.64
47.80
45.85
✓
✓
74.69
46.16
45.12
✓
✓
74.76
48.95
45.97
TABLE XIII: The effectiveness of individual modules on DOTA-v1.0.
Fig. 8: Convergence curves of Deformable DETR-O with CLS, ARS-DETR, OrientedFormer with ResNet50, Swin-T and LSK-T on DIOR-R.
Fig. 9: Epochs of training stage versus accuracy on the DOTA-v1.0 test set.
Method
Frame
Backbone
FPS
Params
FLOPs
AP 50
RoI Transformer [ 41 ]
Two- stage
R50
9.2
55M
253G
74.61
Oriented RCNN [ 4 ]
R50
7.3
41M
225G
75.87
Gliding Vertex [ 6 ]
R101
10.2
41M
225G
75.02
R3Det [ 15 ]
One- stage
R101
6.1
42M
335G
73.79
CFA [ 18 ]
R50
16.6
37M
194G
74.51
SASM [ 8 ]
R50
15.8
37M
194G
74.92
TABLE XIV: Speed, Parameters, FLOPs and accuracy on DOTA-v1.0.
Method
Backbone
Layers
AP 50
AP 75
AP 50:95
Params
FLOPs
Oriented -Former
R50
1
56.21
33.54
33.23
41M
287G
2
60.69
39.86
38.35
41M
297G
3
66.37
44.14
42.35
42M
315G
4
67.28
44.13
42.66
44M
325G
TABLE XV: Comparsions of different layers of Backbone on DIOR-R.
Methods
AP 50
AP 75
OrientedFormer
Fixed Offsets
65.98
43.02
Deformable Offsets
66.66
44.10
Random Offsets
66.70
43.47
Oriented Cross-attention
67.28
44.13
TABLE XVI: Comparison of Different Sampling Methods on DIOR-R.
Fig. 10: Comparison between our method and others on challenging samples. Confidence threshold is set to 0.3. Blue circles denote error results. Yellow circles denote missed results. The 1st line: Large-scale objects, e.g., soccer fields and ground track fields. The 2nd line: Densely packed objects, e.g., planes. The 3rd line: Images with complex backgrounds, e.g., the disturbance of letters. The 4th line: Images under poor environmental conditions.
Fig. 11: Sampling points in cross-attention. Oriented cross-attention (left column) rotates sampling points by angles for alignment. Deformable offsets method (middle column) does not rotate sampling points. Random offsets method (right column) employs random sampling points.
Fig. 12: Center points of learned positonal queries.
Fig. 13: Visualization of oriented cross-attention. For readibility, we draw the center of positional queries and sampling points.
Fig. 14: Visualization results of our method for DOTA, DIOR, and HRSC2016. DOTA contains images depicting extreme weather and poor lighting conditions. Oriented boxes, labels, and confidences are drawn.
School of Computer Science and Technology/School of Artificial Intelligence, the Mine Digitization Engineering Research Center of the Ministry of Education, and Jiangsu Provincial Industrial Technology Engineering Center for Intelligent Sensing and Emergency IoT in Underground Space, China University of Mining and Technology, Xuzhou 221116, China · School of Electrical Engineering and Computer Science, University of Ottawa, Ottawa, ON K1N 6N5, Canada