Authors: Seungwoo Yoon, Dohyun Kang, Eunsue Choi, Sohyun Lee, Seoyeon Kim, Minho Choi, Hyeonsu Heo, Dong-ha Shin, +4 more
Organizations: Department of Computer Science and Engineering, Pohang University of Science and Technology (POSTECH), Pohang 37673, Republic of Korea · Department of Mechanical Engineering, Pohang University of Science and Technology (POSTECH), Pohang 37673, Republic of Korea · Graduate School of Artificial Intelligence, Pohang University of Science and Technology (POSTECH), Pohang 37673, Republic of Korea · Department of Electrical and Computer Engineering, University of Washington, Seattle, 98195, WA, USA · Department of Electrical Engineering, Ulsan National Institute of Science and Technology, Ulsan 44919, Republic of Korea. · Department of Physics, University of Washington, Seattle, 98195, WA, USA · Department of Chemical Engineering, Pohang University of Science and Technology (POSTECH), Pohang 37673, Republic of Korea · Department of Electrical Engineering, Pohang University of Science and Technology (POSTECH), Pohang 37673, Republic of Korea · POSCO-POSTECH-RIST Convergence Research Center for Flat Optics and Metaphotonics, Pohang 37673, Republic of Korea · National Institute of Nanomaterials Technology (NINT), Pohang 37673, Republic of Korea
Obstructions such as raindrops, fences, or dust degrade captured images, especially when mechanical cleaning is infeasible. Conventional solutions to obstructions rely on a bulky compound optics array or computational inpainting, which compromise compactness or fidelity. Metalenses composed of subwavelength meta-atoms promise compact imaging, but simultaneous achievement of broadband and obstruction-free imaging remains a challenge, since a metalens that images distant scenes across a broadband spectrum cannot properly defocus near-depth occlusions. Here, we introduce a learned split-spectrum metalens that enables broadband obstruction-free imaging. Our approach divides the spectrum of each RGB channel into pass and stop bands with multi-band spectral filtering and learns the metalens to focus light from far objects through pass bands, while filtering focused near-depth light through stop bands. This optical signal is further enhanced using a neural network. Our learned split-spectrum metalens achieves broadband and obstruction-free imaging with relative PSNR gains of 32.29% and improves object detection and semantic segmentation accuracies with absolute gains of +13.54% mAP, +48.45% IoU, and +20.35% mIoU over a conventional hyperbolic design. This promises robust obstruction-free sensing and vision for space-constrained systems, such as mobile robots, drones, and endoscopes.
Figures & tables
Fig. 1: De-occluding broadband metalens for obstruction-free broadband imaging. Our de-occluding broadband metalens enables obstruction-free broadband imaging by filtering out the focused light from near depth with a multi-band spectral filter that selectively transmits far-focused light. An LCP generator and an RCP analyzer are included to select the cross-polarized circular component required for the current geometric-phase implementation.
Fig. 2: Depth-wavelength symmetry. a Schematic of the depth-wavelength relationship of diffractive lenses. (Left) An incident wavelength λ longer than the design wavelength λd ( λ>λd ) of the lens causes a phase mismatch, resulting in a focal front shift. (Right) This focal front shift can be compensated for by the spherical phase of an incident wave originating from a near depth, leading to an opposing focal back shift. Here, f is the focal length. Blue and green denote ( λ=λd ) and ( λ>λd ), respectively. b Focal point intensity map for point light sources of wavelength, λ , and depth, z , incident on a hyperbolic ( λd=450 nm) metalens, computed with Rayleigh-Sommerfeld diffraction integral, and overlaid with our depth-wavelength symmetry model.
Fig. 3: Learning de-occluding broadband metalens. a Overview of the metalens optimization pipeline. The metalens design map θ(x,y) is learned for the de-occluding broadband imaging. b Transmission of the multi-band spectral filter, separating each color channel into the pass band Λpassc and the stop band Λstopc , providing extended design space for obstruction-free imaging. c Our compact prototype consists of a metalens, a multi-band spectral filter, a color sensor, and polarization optics for the geometric phase modulation. d x−λ scan result of the learned metalens for far point spread functions (PSFs) and near PSFs in simulation. e Training loss and visual convergence during learning. The sequence of the simulated sensor images (step 50, 200, 3200) illustrates the learning of the metalens. Initial results exhibit significant blur and interference from obstructions, while the final optimized metalens effectively blurs the near-depth obstructions and maintains a sharp focus on the far-depth scene. The source image shown in a and e is from the DIV2K dataset [ 1 ] .
Fig. 4: Characterization of the de-occluding broadband metalens. a Scanning electron microscope (SEM) images of the de-occluding broadband metalens with the top-down view (left) and tilted view (right). b Photograph of the three fabricated metalenses. From left to right: the de-occluding broadband metalens (Ours), the broadband metalens without a de-occluding approach (Broadband), and the conventional hyperbolic-phase metalens (Hyperbolic). c Point spread function (PSF) x−λ scan results normalized to the total intensity. Pass bands are highlighted. Far-depth focused lights of the obstruction-free metalens (Ours) can be transmitted and captured on the sensor, while near-depth focused lights are filtered out, enabling obstruction-free imaging. d 2D PSFs of metalenses for different depths, normalized to their peak intensities. The stop band PSFs of Ours are marked with black boxes. c, d Our far-depth PSFs remain sharp in the pass bands, enabling clear imaging of the originally occluded scene, while the near-depth PSFs are blurry. The hyperbolic metalens suffers from severe chromatic aberrations. The broadband metalens fails to achieve the near-depth blur required for obstruction-free imaging.
Fig. 5: Experimental evaluation of de-occluding broadband imaging. a Near-depth obstruction pattern used for the experiment, placed at znear≃0.045m . Objects and 2D images in the scene are placed at zfar≃0.6m . b Raw sensor captured images under the obstruction for the three metalens designs. Insets provide a magnified view of the regions highlighted by yellow and red boxes. c Ground-truth (GT) unobstructed reference image captured with a compound lens with f=8 mm (two times longer than those of metalenses) for better object plane resolution. d Reconstructed images using the neural network trained for each metalens. e Reconstructed image comparisons of representative printed DIV2K [ 1 ] images under fence and dirt obstructions. f PSNRs measured on the printed DIV2K images [ 1 ] . b, d, e, f Our design suppresses obstructions better than other metalens baselines, achieving superior performance across all conditions.
Fig. 6: Obstruction-free visual perception. a–c Input images captured with each fabricated metalens design acquired by imaging the printed Digital GT, together with the corresponding original digital images (rows: Hyperbolic, Broadband, Ours, Digital GT) under different obstructions (fence, blood drop, dirt) and the corresponding downstream task outputs (Annotation) predicted by computer vision models on three benchmarks and ground-truth label: a UAV-view object detection [ 25 ] on VisDrone [ 57 ] , b Endoscopy polyp segmentation [ 56 ] on Kvasir-SEG [ 23 ] , and c Driving scene semantic segmentation [ 52 ] on Cityscapes [ 13 ] . The bottom row (Digital GT) shows the original digital images and their ground-truth labels for reference. d–f Quantitative performance under unobstructed (purple) and obstructed (green) imaging conditions: d detection mean average precision (mAP) on VisDrone, e intersection-over-union (IoU) on Kvasir-SEG, and f mean IoU (mIoU) on Cityscapes over multiple classes. Across all tasks, our de-occluding broadband metalens exhibits minimal performance degradation under obstructed conditions. Additional real-scene downstream-vision demonstrations are provided in Supplementary Note 17, as well as Supplementary Videos 4 and 5.
Image quality in modern imaging systems emerges from the coupled effects of the sensor, optics, and computational reconstruction. Ultra-thin metalenses offer a path toward substantial miniaturization of optical modules, but practical designs often exhibit pronounced chromatic and field-dependent aberrations that necessitate computational reconstruction. In current metalens pipelines, reconstruction models are commonly trained and selected using distortion-based fidelity objectives, such as PSNR, yet these proxies can be weakly correlated with human preference and downstream utility, reflecting the well-known perception--distortion trade-off. We introduce MetaRanker, a human-in-the-loop active ranking framework that formalizes metalens image quality in terms of semantic interpretability, defined as the degree to which humans can reliably recognize objects and structures in the presence of optical artifacts. MetaRanker combines a probabilistic preference model with uncertainty-aware query selection, and leverages vision--language models to provide lightweight semantic priors. Importantly, these priors are used only to guide the sampling of informative comparisons; human judgments remain the primary supervision signal throughout. Across real-world and synthetic metalens datasets with distinct degradation profiles, MetaRanker produces rankings that align most closely with human assessments, while reducing the number of pairwise annotations required by approximately 80% relative to exhaustive pairwise evaluation. Finally, we show that standard image quality assessment metrics exhibit limited alignment with human interpretability in the metalens domain, positioning MetaRanker as a practical step toward perceptually grounded metalens evaluation and co-design.
Yujin Park, Haejun Chung, Ikbeom Jang
Hanyang University Seoul, Republic of Korea · Hankuk University of Foreign Studies Yongin, Republic of Korea
Hyperspectral 3D imaging enables the capture of dense spectral information and scene geometry but has traditionally been confined to narrow spectral windows, typically the visible range. In this work, we introduce a broadband hyperspectral 3D imaging (BH3D) method to extend this capability across the full visible-near-infrared and short-wavelength infrared (SWIR) spectrum (450-1500 nm). This broad coverage is critical as it captures complementary physical cues: visible wavelengths reveal surface appearance, while SWIR bands provide insight into subsurface properties and material composition. However, realizing BH3D is challenging due to fundamental sensor constraints between visible-spectrum silicon and SWIR-spectrum InGaAs sensors, which necessitate complex multi-spectrograph designs. Here we propose a single-spectrograph BH3D system, using a stereo setup comprising visible and SWIR cameras, that reconstructs dense broadband hyperspectral reflectance together with accurate 3D geometry. Our key idea is to extend dispersed structured light to the broadband regime using a single spectrograph. We model the image formation of broadband dispersed structured light, and estimate hyperspectral reflectance and depth. We validate our approach on diverse real-world scenes, demonstrating accurate reconstruction with a mean spectral angle mapper of 0.13 rad, root mean square error of 0.03, and mean depth error of 4.5 mm. We further demonstrate identifying metameric materials, performing imaging through opaque layers, uncovering hidden features on banknotes, and revealing blood vessels.
Suhyun Shin, Yunseong Moon, Ryota Maeda +3
POSTECH, South Korea · University of Hyogo, Japan · University of Toronto, Canada
Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to ambiguities in metric scale. We introduce metalenses, an emerging class of ultrathin planar optical elements, as a solution to physically encode missing metric depth cues via nanophotonics. In this paper, we bridge the gap between metalens and DFMs to achieve accurate metric monocular depth sensing. In a single monocular shot, our metalens embeds depth-dependent positional shifts into two polarized optical wavefronts. With an input adaptation strategty, we enable direct fine-tuning that aligns a pretrained DFM with the optical signals. To scale the training data, we further develop a comprehensive simulation pipeline that synthesizes metalens responses from RGB-D datasets, incorporating physical factors to minimize the sim-to-real gap. Experiments demonstrate that this approach outperforms both monocular metric depth estimation and depth-from-defocus baselines, showing an effective pathway for accurate monocular metric depth sensing.