Physics-Guided Spectral Distillation for Underwater Image Enhancement on Resource-Constrained Devices
Authors: Yifan Chen, Kai He, Ye Zheng, Jijun Lu, Zhe Sun, Tao Chen
Organizations: College of Future Information Technology, Fudan University, Shanghai 200433, China · Institute of Artificial Intelligence (TeleAI), China Telecom, China · School of Computer Science and Technology, Harbin Institute of Technology, Weihai 264209, China
Underwater image enhancement is crucial for improving visual perception in marine applications. Existing underwater image enhancement studies mainly focus on enhancement quality and visual fidelity, while rarely considering real-time deployment capability, which is essential for resource-constrained underwater robots. To this end, we introduce a physics-guided spectral distillation (PSD) method, which reduces model capacity for real-time applications while maintaining the high performance of underwater image enhancement models. To decompose the outputs of teacher and student models, PSD adopts a multilevel Haar discrete wavelet transform. It transfers low-frequency color and illumination information as well as high-frequency structural details through band-specific objectives. Moreover, the distillation process of PSD is degradation-aware. We estimate degradation-aware weights through a physical head and combine them with ground-truth-guided reliability masks to selectively retain valuable teacher guidance. Experiments on the UIEB, LSUI, and EUVP datasets validate the effectiveness of the proposed method. Furthermore, we demonstrate the benefits of enhanced images for downstream perception tasks, including object detection. Deployment on a self-developed ROV further demonstrates its practical applicability in real-world underwater scenarios.
Figures & tables
Fig. 1: Accuracy–efficiency comparison across Reti-Diff, Restormer, and NAFNet. PSD improves each compact student at unchanged inference cost; circles denote teachers and annotations report parameter counts.
Fig. 2: Spatial and spectral changes under downsampling by factors of 2 , 4 , and 8 , showing progressive high-frequency detail loss.
Fig. 3: Overview of PSD. (a) The frozen teacher and compact student produce restored images that are decomposed into low- and high-frequency bands. Transmission-derived degradation weights ω and ground-truth-guided reliability masks MR regulate band-specific transfer, while the image-space term complements the spectral objectives. (b) The physical head is calibrated using the underwater image-formation model to estimate channel-wise transmission t and global background light B . The teacher and all PSD-specific components are removed at inference.
Backbone
Source
Parameters (M)
GMACs
Param. retained
GMACs retained
Teacher
Student
Teacher
Student
(%)
(%)
Reti-Diff [ 15 ]
ICLR’ 2025
26.13
1.19
175.01
4.32
4.6
2.5
Restormer [ 40 ]
CVPR’ 2022
26.12
0.46
281.98
5.39
1.8
1.9
NAFNet [ 41 ]
ECCV’ 2022
67.89
0.64
126.17
1.52
0.94
1.2
TABLE I: Complexity of the teacher and compressed student models for a 256×256 input. Retention is the student-to-teacher ratio.
Reti-Diff
Source
UIEB
LSUI
Merged EUVP
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
Teacher
–
24.54
0.9341
28.59
0.8782
22.86
0.8191
Task only
–
24.18
0.9263
24.39
0.8680
22.12
0.8166
CTKD [ 37 ]
AAAI’ 2023
23.76 ↓ 0.42
0.9239 ↓ 0.0024
23.40 ↓ 0.99
0.8629 ↓ 0.0051
22.16 ↑ 0.04
0.8202 ↑ 0.0036
DCKD [ 18 ]
AAAI’ 2025
24.72 ↑ 0.54
0.9324 ↑ 0.0061
22.13 ↓ 2.26
0.8474 ↓ 0.0206
22.57 ↑ 0.45
0.8218 ↑ 0.0052
FreeKD+ [ 42 ]
TPAMI’ 2026
23.74 ↓ 0.44
0.9236 ↓ 0.0027
23.06 ↓ 1.33
0.8687 ↑ 0.0007
22.28 ↑ 0.16
0.8156 ↓ 0.0010
TABLE II: Quantitative comparison on UIEB, LSUI, and the merged EUVP dataset across three backbone architectures. Bold and underlined values denote the best and second-best results among compact student variants within each backbone, respectively. Teacher models are references; arrows show changes from the task-only student.
Fig. 4: Qualitative comparison for (a) Reti-Diff, (b) Restormer, and (c) NAFNet. Columns show input, task-only student, teacher, CTKD, DCKD, FreeKD+, PSD, and ground truth.
Fig. 5: Detail comparison of task-only student, PSD student, teacher, and ground truth. Red-box crops in examples (a) and (b) highlight local contrast, color, and structure.
Reti-Diff
Restormer
NAFNet
Variant
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
Task only
24.39
0.8680
27.79
0.8964
25.31
0.8757
(A) w/o PH
24.58 ↑ 0.19
0.8681 ↑ 0.0001
27.78 ↓ 0.01
0.8960 ↓ 0.0004
26.81 ↑ 1.50
0.8840 ↑ 0.0083
(B) w/o LL
24.46 ↑ 0.07
0.8679 ↓ 0.0001
27.83 ↑ 0.04
0.8975 ↑ 0.0011
26.81 ↑ 1.50
0.8858 ↑ 0.0101
(C) w/o HF
24.49 ↑ 0.10
0.8681 ↑ 0.0001
27.99 ↑ 0.20
0.8978 ↑ 0.0014
27.17 ↑ 1.86
0.8861 ↑ 0.0104
(D) w/o Rel.
24.47 ↑ 0.08
0.8668 ↓ 0.0012
27.89 ↑ 0.10
0.8972 ↑ 0.0008
26.78 ↑ 1.47
0.8864 ↑ 0.0107
TABLE III: Leave-one-component-out ablation on LSUI. PH, LL, HF, Rel., and Img. denote physical weighting, low- and high-frequency transfer, spectral reliability gating, and image-space compensation. All components except the named one are retained; task-only uses none. Arrows indicate performance gains over the task-only baseline.
Fig. 6: Qualitative ablation of PSD on (a) Reti-Diff, (b) Restormer, and (c) NAFNet. From left to right, (A)–(E) denote w/o PH, w/o LL, w/o HF, w/o Rel., and w/o Img., matching the component order in Table III . Full denotes the complete PSD model. Removing individual components produces complementary degradations in color, illumination, contrast, or local structure.
AP50
AP50:95
Input domain
Sea cucumber
Sea urchin
Scallop
mAP
Sea cucumber
Sea urchin
Scallop
mAP
Raw
50.18
87.56
54.46
64.06
21.04
38.01
20.85
26.63
Reti-Diff teacher
58.80
88.97
51.66
66.47
22.85
37.14
20.43
26.80
Reti-Diff student (PSD)
55.13
88.86
54.97
66.32
21.57
40.23
21.19
27.66
Restormer teacher
54.90
89.21
62.33
68.81
22.74
40.57
25.92
29.74
Restormer student (PSD)
58.72
88.86
54.27
67.28
23.85
37.34
21.65
27.61
TABLE IV: YOLOv9s object detection on UDD. Each detector is trained and tested on the corresponding raw or enhanced image domain. Values are percentages; higher is better. The best result in each column is bold.
Fig. 7: YOLOv9s detections on UDD. Columns: Raw; Reti-Diff teacher/student; Restormer teacher/student; NAFNet teacher/student; ground truth. All students use PSD. Detectors are trained and tested in their corresponding domains.
Fig. 8: The self-developed ROV used for field data collection in Dushu Lake. Front and top views show the compact vehicle configuration and the locations of the onboard GoPro and stereo camera.
Backbone
Teacher (FPS)
Student (FPS)
Speedup ( × )
Reti-Diff
5.0
36.4
7.28
Restormer
4.3
35.7
8.30
NAFNet
25.6
84.6
3.30
TABLE V: End-to-end inference speed on the self-developed ROV platform using a fixed 256×256 input. Higher is better.
Fig. 9: SIFT feature matching on images collected by the ROV. The left half shows matches between raw frame pairs (yellow), and the right half shows matches between the corresponding NAFNet-student-enhanced pairs (green). The numbers below the panels denote the retained matches for each pair.
Metric
Raw
Enhanced
Gain (%)
UIQM ↑
1.7902
2.2276
24.43
UCIQE ↑
0.2184
0.2830
29.57
TABLE VI: No-reference quality scores on the images collected in Dushu Lake. Higher is better.
Glasgow College, UESTC, Chengdu, China · University of Glasgow, Glasgow, UK · School of Information and Communication Engineering, UESTC, Chengdu, China
College of Future Information Technology, Fudan University, Shanghai 200433, China · Institute of Artificial Intelligence (TeleAI), China Telecom, China · Faculty of Robot Science and Engineering, Northeastern University, Shenyang 110819, China +1