LoopLUT: 3D Lookup Tables with Progressive Region Refinement for Real-Time 4K Image Enhancement
Authors: Yang Ye, Jiajun Ma, Chen Wu, Wei Wang, Dianjie Lu, Guijuan Zhang, Linwei Fan, Zhuoran Zheng
Organizations: Universiti Sains Malaysia · National University of Defense Technology · Sun Yat-sen University · Shandong Normal University · Shandong University of Finance and Economics
Color enhancement of 4K images must meet a quality target under a tight compute budget. Three-dimensional lookup tables (3D LUTs) dominate real-time enhancement because they decide at low resolution and apply a per-pixel lookup at full resolution. A single global LUT, however, is spatially invariant, so an underexposed shadow and a well-exposed region that share a pixel value receive identical corrections. Spatially heterogeneous demands cannot be expressed by such a mapping. We propose LoopLUT, a region-cascaded 3D LUT with progressive refinement. A global LUT performs the overall correction, followed by K-1 loop iterations. In each iteration a gating head predicts at low resolution the region that still needs correction, then builds a residual LUT from the color statistics of that region alone. The cascaded gates form a partition of unity, so the output is a per-pixel convex combination of the K lookup results. Fusion is therefore performed by the gates themselves, with no separate fusion module and no interpolation error accumulating across rounds. The decision stage runs at a fixed 256x256 resolution, independent of output resolution, so a 4K image costs only K pure lookups. Extensive experiments across four benchmarks show that LoopLUT improves PSNR by up to 2.81 dB over the strongest prior method, while keeping real-time throughput at 4K. The same decomposition also generalizes well to underwater enhancement datasets.
Figures & tables
Figure 1: LoopLUT divides an image among its correction rounds. (a) An LCDP test image. (b–d) Partition weights from the trained gates of LoopLUT ( K=3 ), shown without contrast stretching; color saturation encodes the weight. Ak(x) is the share of pixel x taken from the lookup of round k . The means (0.65, 0.20 and 0.15) show that the global round covers most of the image and later rounds refine ever smaller regions. Each weight is what the cascade passes to round k minus what round k passes on, Ak=Ck(1−mk+1) , where Ck is the product of the gates up to round k . The weights are therefore non-negative and sum to one at every pixel, so the output is a convex blend of the K lookups. (e) Profile along the dashed scanline, confirming the unit sum and smooth transitions.
Figure 2: Overall architecture of LoopLUT ( Mn , Cn and Tn in the figure are mk , Ck and Tk in the text). The global round (gray) passes the 256×256 input once through the Color UNet to give the global feature Fg (VectorG) and the decoder feature map D (Feature map), which an MLP turns into the global LUT T1 by mixing the basis LUTs B1,…,B5 . Each region round (blue, ×K−1 ) area-pools its inputs to the s×s gate bottleneck, predicts the gate Mn , multiplies it into the cumulative gate Cn , and generates Tn from masked statistics inside Cn . At execution the K LUTs are applied to the original image and combined per pixel by the partition weights, so the gates are the fusion. The decision stage is fixed at low resolution throughout.
LCDP
MIT-Adobe FiveK
Mobile-Spec
Method
Venue
PSNR ↑
SSIM ↑
LPIPS ↓
CIEDE ↓
PSNR ↑
SSIM ↑
LPIPS ↓
CIEDE ↓
PSNR ↑
SSIM ↑
LPIPS ↓
CIEDE ↓
HDRNet
SIGGRAPH’17
21.75
0.8265
0.1860
10.98
20.57
0.7251
0.0960
9.10
20.34
0.7911
0.1652
6.57
RUAS
CVPR’21
18.25
0.7568
0.2065
9.11
19.29
0.7095
0.1393
9.86
12.22
0.7470
0.1833
17.26
LCDPNet
ECCV’22
23.29
0.8423
0.1587
8.43
15.13
0.6029
0.2437
17.17
25.25
0.9097
0.1404
5.21
AdaInt
CVPR’22
20.93
0.7856
0.2258
9.42
18.76
0.7235
0.1474
12.21
26.98
0.9264
0.0850
4.02
LightenDiffusion
ECCV’24
19.04
0.7743
0.1334
9.33
18.35
0.7050
0.1542
12.35
23.23
0.8383
0.2759
7.31
Table 1: Quantitative comparison with state-of-the-art methods on LCDP, MIT-Adobe FiveK and Mobile-Spec. In each column the best is bold and the second best is underlined; ↑ / ↓ indicates that higher/lower is better.
Figure 3: Qualitative comparison on the LCDP test set. Each panel shows the full output above and, below, the magnified red and green crops, placed on the two middle paintings and on the steel frame at the top; the red box examines midtone color reproduction and the green box detail recovery in dark structure. All methods use the same crops.
Figure 4: Qualitative comparison on the MIT-Adobe FiveK test set. Layout as in Figure 3 ; the red crop covers the face of a stone head and the green crop the weathered stone of its crown, examining the light-to-dark transition across the face and texture preservation on the highlight side.
Figure 5: Qualitative comparison on the Mobile-Spec test set. Layout as in Figure 3 ; the red crop covers the sunlit facade and the green crop the backlit region at lower left, examining color reproduction under warm illumination and tonal separation in the backlit building and tree canopy.
Figure 7
Figure 7: Ablation 3: progressive refinement on an LCDP test example. PSNR rises 19.16 (input) → 20.56 → 23.74 → 24.04 dB, and each row’s “Current input” is the previous row’s “Round output”. Round 1 is the global round, so its overlay is entirely green. In the partition overlay, green marks pixels fixed after round 1 ( A1 ), red those entering round 2 ( A2 ) and yellow round 3 ( A3 ); “Hard selection” is a thresholded view of the gate, shown for illustration only.
Resolution
Pixels
GFLOPs
Latency
Device memory
(MP)
Decide
Exec
Total
Decide
Total
Through-
Peak
Traffic
Band-
(ms)
(ms)
put (FPS)
(MB)
(MB)
width (GB/s)
1620×1080 (LCDP native)
1.75
6.135
0.362
6.497
1.84
5.10
196
93
774
159
1920×1080 (FHD)
2.07
6.135
0.429
6.564
1.84
5.67
176
104
880
163
3840×2160 (4K)
8.29
6.135
1.717
7.852
1.84
16.23
62
296
2914
188
Table 3: Resolution scaling and memory behavior of LoopLUT. The decision stage is constant at 6.135 GFLOPs and 1.84 ms at every resolution ; only the pure-lookup execution stage grows with pixel count. Traffic is the per-frame device-memory read + write volume of the current unfused implementation, and Bandwidth is that volume divided by the measured latency.
Table 4: LSUI underwater benchmark. Best in bold, second underlined; compared methods quoted from public reports.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Venue
PSNR ↑
SSIM ↑
LPIPS ↓
CIEDE ↓
HDRNet
SIGGRAPH’17
20.89
0.7809
0.1491
8.88
RUAS
CVPR’21
16.59
0.7378
0.1764
12.08
LCDPNet
ECCV’22
21.22
0.7850
0.1809
10.27
AdaInt
CVPR’22
22.22
0.8118
0.1527
8.55
LightenDiffusion
ECCV’24
20.21
0.7725
0.1878
9.66
CSEC
CVPR’24
20.12
0.7748
0.1962
11.62
Appendix
Table 5: Average over LCDP, MIT-Adobe FiveK and Mobile-Spec of the metrics in Table 1 . In each column the best is bold and the second best is underlined.
Figure 9: The averages of Table 5 , with methods sorted from best to worst in each panel; the arrow marks the gap between LoopLUT and the runner-up.
Figure 10: Visual comparison on four LSUI test images, one image per row. The compared methods are those of Table 8 .
Figure 11: Real-world smartphone photographs. Top: an underexposed indoor scene; middle: a night scene with blown-out street lamps (same crop for all methods). Bottom: NIQE against latency (left) and averages over the two photographs (right; best in bold).
Figure 12: Mobile deployment. LoopLUT runs on an Android phone through the ONNX Runtime CPU backend, with three tabs corresponding to (a) input, (b) enhanced output and (c) partition assignment. The colors in (c) are the partition weights of Equation ( 1 ): green for the global round A1 , yellow for A2 and red for A3 , displayed after per-image contrast stretching; the three thumbnails below can be tapped to enlarge. The bottom of the interface shows the measured readout for that particular run : Decide 129 ms, Render 52 ms, total 198 ms. On-device total time fluctuates between roughly 100 and 200 ms depending on the input photograph. Decide corresponds to the fixed-resolution decision stage and Render to the execution stage that grows with pixel count, matching the split in Table 3 .
Underwater image enhancement is challenged by spatially non-uniform, wavelength-dependent attenuation. Propagation distance and wavelength govern this degradation, while YCbCr separates luminance from chrominance for restoration. We propose DY-LUT, a depth-aware YCbCr lookup-table framework for real-time enhancement. A dual-branch encoder predicts image-level fusion weights and a joint pair of pixel-wise degradation indices from image and depth features. These quantities condition learnable 4D LUTs, followed by lightweight local refinement. DY-LUT preserves traditional LUT efficiency while enabling depth-conditioned, spatially adaptive restoration. With externally supplied depth, its 3.56M-parameter enhancement network achieves competitive quality on UIEB-90 and LSUI and runs 9--304× faster than representative high-capacity baselines. Adaptive inference further maintains real-time performance (∼7 ms) for 4K UIQAD images. DY-LUT also benefits downstream detection and feature matching. Ablations show that YCbCr is a more effective basis than RGB for depth-conditioned lookup, while the jointly learned indices further improve adaptive querying. These results provide a physically grounded route to efficient UIE on practical platforms.
Cunhao Zhu, Xiangtao Kong, Dongliang Xu +3
1Shandong University · 2The Hong Kong Polytechnic University · 3Mohamed bin Zayed University of Artificial Intelligence
Photographic color editing is inherently personal: the same image can appear too warm, too muted, or already satisfactory to different users. Most lookup table (LUT) and reference-guided methods target a specified appearance rather than model persistent preferences from repeated user choices. To address this gap, we introduce PrefLUT, a reusable and refinable user-preference modeling framework for deployable 3D LUTs, encoding ordered preferred/non-preferred image pairs into a lightweight Reusable User Profile that is reused across queries and refined using additional user preference pairs, without per-user optimization. A Query-Conditioned LUT Predictor combines this profile with each image to predict a LUT latent vector and edit strength. An Identity-Residual LUT Decoder and Edit-Strength Controller then produce an exportable 3D LUT. Experiments on three datasets demonstrate effective personalized editing and general-purpose enhancement. Each quantized profile requires only 260 bytes, and editing takes 1.365 ms/image on an RTX 5090 GPU. We also introduce the Preference-Conditioning Verification Protocol (PCVP), an evaluation protocol to verify whether personalized image edits depend on user preferences and the query image through controlled changes to user profiles, preference orders, pair correspondences, and query images.
Chuanzhi Xu, Langyi Chen, Chengkun Yue +5
The University of Sydney · Indiana University · The University of Hong Kong
Lookup table (LUT)-based image denoising methods have attracted increasing attention due to their high efficiency and hardware-friendly properties. However, existing RGB-LUT approaches require three identical LUTs to process RGB channels in parallel, resulting in large on-chip SRAM consumption. A simple alternative is to apply LUT processing only to the luminance (Y) channel in the YUV color space to reduce memory usage. However, this naive strategy leads to degraded restoration quality, since ignoring the chrominance (UV) channels introduces color distortion and residual artifacts. In this work, we propose Hybrid-LUT, a YUV-based asymmetric channel-processing framework that combines LUT and filtering in a unified design. Specifically, a multi-band LUT branch with pixel-level weight fusion is applied to the Y channel to recover fine textures, while lightweight filtering is used for the UV channels to maintain color consistency. This design reduces LUT storage by two-thirds compared with RGB-LUT methods while maintaining the same runtime throughput. Extensive experiments show that Hybrid-LUT achieves state-of-the-art (SOTA) performance across multiple benchmarks with only 421 KB of storage. In particular, our method surpasses existing LUT-based denoising approaches by at least 0.63 dB CPSNR on real-world datasets, demonstrating its effectiveness for image denoising on resource-constrained edge devices. The project is available at https://github.com/Ai-ZL/Hybrid-LUT .