Physics-Guided Spectral Distillation for Underwater Image Enhancement on Resource-Constrained Devices
Authors: Yifan Chen, Kai He, Ye Zheng, Jijun Lu, Zhe Sun, Tao Chen
Organizations: College of Future Information Technology, Fudan University, Shanghai 200433, China · Institute of Artificial Intelligence (TeleAI), China Telecom, China · School of Computer Science and Technology, Harbin Institute of Technology, Weihai 264209, China
Underwater image enhancement is crucial for improving visual perception in marine applications. Existing underwater image enhancement studies mainly focus on enhancement quality and visual fidelity, while rarely considering real-time deployment capability, which is essential for resource-constrained underwater robots. To this end, we introduce a physics-guided spectral distillation (PSD) method, which reduces model capacity for real-time applications while maintaining the high performance of underwater image enhancement models. To decompose the outputs of teacher and student models, PSD adopts a multilevel Haar discrete wavelet transform. It transfers low-frequency color and illumination information as well as high-frequency structural details through band-specific objectives. Moreover, the distillation process of PSD is degradation-aware. We estimate degradation-aware weights through a physical head and combine them with ground-truth-guided reliability masks to selectively retain valuable teacher guidance. Experiments on the UIEB, LSUI, and EUVP datasets validate the effectiveness of the proposed method. Furthermore, we demonstrate the benefits of enhanced images for downstream perception tasks, including object detection. Deployment on a self-developed ROV further demonstrates its practical applicability in real-world underwater scenarios.
Figures & tables
Fig. 1: Accuracy–efficiency comparison across Reti-Diff, Restormer, and NAFNet. PSD improves each compact student at unchanged inference cost; circles denote teachers and annotations report parameter counts.
Fig. 2: Spatial and spectral changes under downsampling by factors of 2 , 4 , and 8 , showing progressive high-frequency detail loss.
Fig. 3: Overview of PSD. (a) The frozen teacher and compact student produce restored images that are decomposed into low- and high-frequency bands. Transmission-derived degradation weights ω and ground-truth-guided reliability masks MR regulate band-specific transfer, while the image-space term complements the spectral objectives. (b) The physical head is calibrated using the underwater image-formation model to estimate channel-wise transmission t and global background light B . The teacher and all PSD-specific components are removed at inference.
Backbone
Source
Parameters (M)
GMACs
Param. retained
GMACs retained
Teacher
Student
Teacher
Student
(%)
(%)
Reti-Diff [ 15 ]
ICLR’ 2025
26.13
1.19
175.01
4.32
4.6
2.5
Restormer [ 40 ]
CVPR’ 2022
26.12
0.46
281.98
5.39
1.8
1.9
NAFNet [ 41 ]
ECCV’ 2022
67.89
0.64
126.17
1.52
0.94
1.2
TABLE I: Complexity of the teacher and compressed student models for a 256×256 input. Retention is the student-to-teacher ratio.
Reti-Diff
Source
UIEB
LSUI
Merged EUVP
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
Teacher
–
24.54
0.9341
28.59
0.8782
22.86
0.8191
Task only
–
24.18
0.9263
24.39
0.8680
22.12
0.8166
CTKD [ 37 ]
AAAI’ 2023
23.76 ↓ 0.42
0.9239 ↓ 0.0024
23.40 ↓ 0.99
0.8629 ↓ 0.0051
22.16 ↑ 0.04
0.8202 ↑ 0.0036
DCKD [ 18 ]
AAAI’ 2025
24.72 ↑ 0.54
0.9324 ↑ 0.0061
22.13 ↓ 2.26
0.8474 ↓ 0.0206
22.57 ↑ 0.45
0.8218 ↑ 0.0052
FreeKD+ [ 42 ]
TPAMI’ 2026
23.74 ↓ 0.44
0.9236 ↓ 0.0027
23.06 ↓ 1.33
0.8687 ↑ 0.0007
22.28 ↑ 0.16
0.8156 ↓ 0.0010
TABLE II: Quantitative comparison on UIEB, LSUI, and the merged EUVP dataset across three backbone architectures. Bold and underlined values denote the best and second-best results among compact student variants within each backbone, respectively. Teacher models are references; arrows show changes from the task-only student.
Fig. 4: Qualitative comparison for (a) Reti-Diff, (b) Restormer, and (c) NAFNet. Columns show input, task-only student, teacher, CTKD, DCKD, FreeKD+, PSD, and ground truth.
Fig. 5: Detail comparison of task-only student, PSD student, teacher, and ground truth. Red-box crops in examples (a) and (b) highlight local contrast, color, and structure.
Reti-Diff
Restormer
NAFNet
Variant
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
Task only
24.39
0.8680
27.79
0.8964
25.31
0.8757
(A) w/o PH
24.58 ↑ 0.19
0.8681 ↑ 0.0001
27.78 ↓ 0.01
0.8960 ↓ 0.0004
26.81 ↑ 1.50
0.8840 ↑ 0.0083
(B) w/o LL
24.46 ↑ 0.07
0.8679 ↓ 0.0001
27.83 ↑ 0.04
0.8975 ↑ 0.0011
26.81 ↑ 1.50
0.8858 ↑ 0.0101
(C) w/o HF
24.49 ↑ 0.10
0.8681 ↑ 0.0001
27.99 ↑ 0.20
0.8978 ↑ 0.0014
27.17 ↑ 1.86
0.8861 ↑ 0.0104
(D) w/o Rel.
24.47 ↑ 0.08
0.8668 ↓ 0.0012
27.89 ↑ 0.10
0.8972 ↑ 0.0008
26.78 ↑ 1.47
0.8864 ↑ 0.0107
TABLE III: Leave-one-component-out ablation on LSUI. PH, LL, HF, Rel., and Img. denote physical weighting, low- and high-frequency transfer, spectral reliability gating, and image-space compensation. All components except the named one are retained; task-only uses none. Arrows indicate performance gains over the task-only baseline.
Fig. 6: Qualitative ablation of PSD on (a) Reti-Diff, (b) Restormer, and (c) NAFNet. From left to right, (A)–(E) denote w/o PH, w/o LL, w/o HF, w/o Rel., and w/o Img., matching the component order in Table III . Full denotes the complete PSD model. Removing individual components produces complementary degradations in color, illumination, contrast, or local structure.
AP50
AP50:95
Input domain
Sea cucumber
Sea urchin
Scallop
mAP
Sea cucumber
Sea urchin
Scallop
mAP
Raw
50.18
87.56
54.46
64.06
21.04
38.01
20.85
26.63
Reti-Diff teacher
58.80
88.97
51.66
66.47
22.85
37.14
20.43
26.80
Reti-Diff student (PSD)
55.13
88.86
54.97
66.32
21.57
40.23
21.19
27.66
Restormer teacher
54.90
89.21
62.33
68.81
22.74
40.57
25.92
29.74
Restormer student (PSD)
58.72
88.86
54.27
67.28
23.85
37.34
21.65
27.61
TABLE IV: YOLOv9s object detection on UDD. Each detector is trained and tested on the corresponding raw or enhanced image domain. Values are percentages; higher is better. The best result in each column is bold.
Fig. 7: YOLOv9s detections on UDD. Columns: Raw; Reti-Diff teacher/student; Restormer teacher/student; NAFNet teacher/student; ground truth. All students use PSD. Detectors are trained and tested in their corresponding domains.
Fig. 8: The self-developed ROV used for field data collection in Dushu Lake. Front and top views show the compact vehicle configuration and the locations of the onboard GoPro and stereo camera.
Backbone
Teacher (FPS)
Student (FPS)
Speedup ( × )
Reti-Diff
5.0
36.4
7.28
Restormer
4.3
35.7
8.30
NAFNet
25.6
84.6
3.30
TABLE V: End-to-end inference speed on the self-developed ROV platform using a fixed 256×256 input. Higher is better.
Fig. 9: SIFT feature matching on images collected by the ROV. The left half shows matches between raw frame pairs (yellow), and the right half shows matches between the corresponding NAFNet-student-enhanced pairs (green). The numbers below the panels denote the retained matches for each pair.
Metric
Raw
Enhanced
Gain (%)
UIQM ↑
1.7902
2.2276
24.43
UCIQE ↑
0.2184
0.2830
29.57
TABLE VI: No-reference quality scores on the images collected in Dushu Lake. Higher is better.
Real-time underwater image enhancement (UIE) is crucial for mobile underwater photography and autonomous robotic systems, where practical deployment typically requires low latency and compact models under constrained computational resources. Recent ultra-lightweight CNNs based on structural re-parameterization meet these constraints but operate purely in the spatial domain, ignoring the frequency-sensitive nature of underwater degradation. To address this, we propose a lightweight UIE framework that integrates two key components: a Multi-Branch Reparameterizable Convolution with Fixed DCT Priors (MBRConv-DCT) that injects structured directional frequency priors during training, and a Frequency-Guided Dual-Path Attention (FGDPA) module that fuses spatial and spectral cues via a dual-path design for adaptive feature modulation. Both components are fully compatible with structural re-parameterization: the convolution branch introduces zero additional inference cost after re-parameterization, while the attention module incurs only a minimal computational overhead. Experiments show our model achieves state-of-the-art performance with only 4.23K parameters and 600+ FPS, outperforming much larger methods in both quantitative metrics and visual quality. Code is available at https://github.com/LethyZhang/FGDPA.
Leshen Zhang, Ao Li, Ce Zhu
Glasgow College, UESTC, Chengdu, China · University of Glasgow, Glasgow, UK · School of Information and Communication Engineering, UESTC, Chengdu, China
Underwater image enhancement (UIE) aims to recover clear images from observations affected by wavelength-dependent absorption, scattering, and spatially nonuniform degradation. Although existing generative methods can handle complex degradations, severe information loss may lead to semantic drift in the restored results. To address this issue, we propose RPL-UIE, a two-stage teacher--student framework for reliable prior learning. In the teacher stage, the network learns reliable and complementary spatial priors characterizing appearance and photometric properties from paired degraded and reference images. In the student stage, the network takes only degraded images as input and learns to emulate the teacher's prior extraction capability, thereby providing more reliable restoration guidance for the enhancement process without requiring reference images at inference. To reduce the prior-learning discrepancy between the teacher and student models, we further develop Residual Prior Refinement Diffusion (RPRD) and Frequency-Aware Prior Residual Calibration (FPRC). RPRD uses the coarse priors as anchors and progressively predicts the necessary corrections in the residual space. FPRC retains stable low-frequency residual components and selectively modulates high-frequency detail residuals, producing calibrated priors to support high-quality reconstruction. Experiments on multiple UIE benchmarks demonstrate competitive restoration performance. Downstream underwater object detection and instance segmentation experiments further demonstrate the improved utility of enhanced images for visual perception, while tests on real-world data captured by a remotely operated vehicle (ROV) support the practical applicability of RPL-UIE.
Yifan Chen, Jiaming Liu, Ye Zheng +2
College of Future Information Technology, Fudan University, Shanghai 200433, China · Institute of Artificial Intelligence (TeleAI), China Telecom, China · Faculty of Robot Science and Engineering, Northeastern University, Shenyang 110819, China +1
Underwater images often suffer from severe degradation, such as color distortion, low contrast, and blurred details, due to light absorption and scattering in water. While learning-based methods like CNNs and Transformers have shown promise, they face critical limitations: CNNs struggle to model the long-range dependencies needed for non-uniform degradation, and Transformers incur quadratic computational complexity, making them inefficient for high-resolution images. To address these challenges, we propose Hero-Mamba, a novel Mamba-based network that achieves efficient dual-domain learning for underwater image enhancement. Our approach uniquely processes information from both the spatial domain (RGB image) and the spectral domain (FFT components) in parallel. This dual-domain input allows the network to decouple degradation factors, separating color/brightness information from texture/noise. The core of our network utilizes Mamba-based SS2D blocks to capture global receptive fields and long-range dependencies with linear complexity, overcoming the limitations of both CNNs and Transformers. Furthermore, we introduce a ColorFusion block, guided by a background light prior, to restore color information with high fidelity. Extensive experiments on the LSUI and UIEB benchmark datasets demonstrate that Hero-Mamba outperforms state-of-the-art methods. Notably, our model achieves a PSNR of 25.802 and an SSIM of 0.913 on LSUI, validating its superior performance and generalization capabilities.
Tejeswar Pokuri, Shivarth Rai
Manipal Institute of Technology, Manipal Academy of Higher Education, Manipal - 576104, India