cs.CVAug 26, 2026

RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

Authors: Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang, Ye Li

Organizations: National Key Laboratory of Autonomous Marine Vehicle Technology, Harbin Engineering University · Department of Aeronautical and Aviation Engineering, Hong Kong Polytechnic University

Abstract

Underwater unimodal object detection faces many challenges in sensor imaging, such as optical images limited by underwater noise and visible distance, and sonar images limited by less object structural information. While, optical images have rich object structural information, and sonar images are less affected by underwater noise and have a longer visible distance. Optical (RGB modality) and sonar (Sonar modality) images have complementary information underwater. In this paper, we create an RGB-Sonar multimodal object detection dataset, \textbf{R}GB-\textbf{S}onar \textbf{Fusion} (RSFusion) and propose evaluation metrics for the benchmark. And we propose the \textbf{R}GB-\textbf{S}onar \textbf{Fusion} \textbf{Det}ector (RSFusionDet) with a new RGB-Sonar multimodal object detection result expression for RGB-Sonar multimodal object detection. We analyze the features of RGB and Sonar modal information, and design a Cross-Attention Fusion (CAFusion) module to fuse RGB-Sonar spatial misalignment features and Object Matching Head (OMHead) with Loss (OMLoss) to match identical objects in RGB-Sonar modalities. Our RSFusionDet achieves 76.4/48.6 AP (RGB/Sonar) for object detection and 83.4 F1-Scorematch\text{F1-Score}_{match} for object matching, on RSFusion, which outperforms other object detection models. Compared with the DINO baseline, our method improves by 0.7/1.4 AP (RGB/Sonar) while simultaneously providing reliable cross-modal object matching. The code and datasets are publicly available at https://github.com/LEFTeyex/RSFusionDet.

Explore similar work

CardsList
  1. A Sonar-Visual Dataset for Cross-Modal Underwater Robot Perception

    May 31, 2026Weitung Chen, Phil Tinn, Per Gunnar Auran +2Cross-Modal LearningUnderwater Robotics

  2. uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception

    Aug 28, 2026Trung Tien Dong, Zhenqi Wu, Aditya Penumarti +4Underwater Robotics3D Scene Understanding