Seeing Through the Glare: A Multi-Source Benchmark and Ocular-Adaptive Pixel MeanFlow for Eyeglass Reflection Removal
Authors: Tao Liu, Youwei Pang, Kailai Zhou, Jiaming Zuo, Hanqi Liu, Wei Ji, Peng-Tao Jiang, Xiaofeng Liu, +2 more
Organizations: AI4X Team, Nanyang Technological University, Singapore · X3000 Inspection Co., Ltd. · Tsinghua University, China · Yale University, USA · vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.
Eyeglass reflection removal is important across smartphone imaging, video conferencing, and other face-centric visual applications. The task is challenging because reflections range from mild photometric contamination to severe ocular occlusion, requiring selective correction and plausible reconstruction without altering identity or natural appearance. Existing datasets cover limited reflection conditions, constraining generalization to complex real-world scenes and systematic evaluation. We introduce \textbf{OcuBench}, a multi-source benchmark comprising 10,280 controllable synthetic pairs, 732 real-input pseudo-pairs, and 458 independent real-world test images, supporting both paired evaluation and assessment beyond generated supervision. We further propose \textbf{OcuFlow}, an ocular-adaptive pixel MeanFlow (pMF) framework for efficient, detail-preserving restoration. It combines geometry-adaptive representation with one-step pMF to focus reconstruction on reflection-obscured ocular regions, together with native-resolution frequency-preserving synthesis to retain reliable observed details. Experiments across diverse reflection conditions demonstrate that OcuFlow achieves consistent advantages in reflection removal quality, ocular fidelity, and efficiency. In a blind user study, it receives 67.32% of selections, 6.2× the next-best share. Both the code and dataset will be released.
Figures & tables
Figure 1 : Eyeglass reflection removal on synthetic samples. Dashed lines separate input (Reflection) and restored (Removal) regions, with complementary comparisons across both rows.
Dataset
Year
Syn.
Real
Real Test
Scale
Access
EwE [ 3 ]
2017
✓
–
–
540
Legacy (Link Unavailable)
Watanabe et al. [ 10 ]
2021
–
✓
–
1,639
No Public Link
De-Glared [ 9 ]
2023
✓
–
–
2,541
No Public Link
ReyeR [ 1 ]
2024
–
✓
–
13,610
No Public Link
ReyeR+ [ 2 ]
2025
–
✓
–
14,328
No Public Link
OcuBench
2026
✓
✓
✓
10,280 / 732 / 458
Public / Verified
Table 1 : Comparison of representative datasets for eyeglass reflection removal. “Syn.” denotes synthetic supervision, “Real” denotes real reflected inputs, and “Real Test” denotes an independent real-world evaluation set.
Figure 2 : Construction of OcuBench. Ocu-S is generated from hierarchically sampled subject, imaging, and reflection factors. Ocu-P retains original reflected images and generates only clean targets. Ocu-R combines captured reference pairs and unpaired images from unconstrained scenes.
Figure 3 : Dataset analysis of OcuBench. (a) Ocu-S reflection configurations. (b) Ocu-S imaging configurations. (c) Empirical cumulative distributions of image properties across OcuBench subsets.
Figure 4 : Overview of OcuFlow. Ocular priors guide allocation of 256 tokens. Geometry, reflected observations, and coarse context condition one-step patch restoration. Frequency-preserving synthesis blends predictions with native details within soft lens support.
Figure 5 : Real-world evaluation on OcuBench. (a) NR-IQA results on 458 real-world images, where the arrow indicates the direction of better performance. (b) Blind user preference comparison; error bars denote 95% cluster-bootstrap confidence intervals.
Table 5 : Primary high-level factor space used for Ocu-S generation.
Figure 7 : Representative examples of mask-aware edge-consistency filtering. Each row shows a clean candidate, its similarity-aligned reflected counterpart, and the corresponding structural-edge overlay. The examples include one accepted pair and three typical failure modes: facial/ocular drift, eyeglass-frame deformation, and unintended edits beyond the lens-reflection region.
Figure 8 : Representative examples of real reflected-input screening. The top row illustrates SoF samples under CLIP-based screening, including a genuine lens reflection and typical negatives such as clear lenses, frame highlights, and sunglasses. The bottom row illustrates Web samples under FaceMesh-guided local screening, including a retained lens-glare example and representative false positives caused by eye whites, frame/specular highlights, and dominant ambient brightness. Cyan boxes indicate the estimated lens regions used for local glare analysis.
Source
Initial
Automatic
Manual
Final
SoF
2,662
400
199
199
Web
11,520
1,740
533
533
Total
14,182
2,140
732
732
Appendix
Table 6 : Construction of the real reflected-input subset. Automatic screening uses CLIP ranking for SoF and FaceMesh-guided reflection screening for Web images.
Figure 9 : Image-level characteristics of the independently collected real-world test set. Smoothed distributions are shown separately for the 157 paired-reference reflected inputs and the 301 unpaired in-the-wild inputs in terms of (a) resolution, (b) aspect ratio, (c) luminance, (d) contrast, (e) high-frequency detail, and (f) color saturation. Dashed vertical lines indicate the median of each subset.
Figure 10 : Construction of the ocular-prior-guided adaptive representation. Face and iris landmarks extracted from the reflected input are used to construct geometry-derived ocular supports, which define the spatial importance field and structural constraints. An exact-budget heap solver allocates N=256 quadtree leaves across four spatial scales, concentrating fine supports on ocular and lens regions while retaining coarse context elsewhere. The resulting leaf supports are mapped back to the native-resolution image for canonical patch sampling, while leaf geometry and pooled semantic priors provide conditioning for subsequent restoration.
Figure 11 : Overview of native-resolution frequency-preserving synthesis. Projected patch predictions form absolute and residual reconstructions. Spatial maps α and β mix L(Iabs) with L(Ires) and H(Ig) with H(Iabs) , respectively. Final lens-supported compositing preserves the observed input outside the spectacle region.
Metric
Region / Purpose
Better
Mask-PSNR / SSIM
Eyeglass region; local restoration fidelity
↑
Mixed-PSNR / SSIM
Eyeglass and non-eyeglass regions; restoration–preservation balance
Figure 12 : Downstream iris-localization evaluation. (a) Both-eyes localization success rate under different iris-circle IoU thresholds. (b) Distribution of the minimum left/right iris-circle IoU for each method. Glare denotes the original reflected input without reflection removal, while OcuFlow is highlighted in red.
Figure 13 : Additional qualitative comparisons on held-out paired OcuBench examples. Enlarged ocular regions are shown below the corresponding full-image comparisons to highlight residual reflections, iris and eyelid structure, and local detail preservation.
Figure 14 : Additional qualitative comparisons on held-out paired OcuBench examples. Methods are shown in the same order as in Fig. 13 .
Figure 15 : Further qualitative comparisons on held-out paired OcuBench examples, covering additional subjects, reflection patterns, and ocular appearances.
Figure 16 : Ocular-prior-guided selective restoration. In the upper example, the eyeglasses are displaced away from the visible eyes, and OcuFlow leaves the lens appearance largely unchanged. In the lower example, a strong reflection overlaps the eye; OcuFlow applies stronger correction to the ocularly relevant region while making smaller changes to surrounding lens reflections.
Figure 17 : Failure case on sunglasses. OcuFlow reconstructs visible eye content behind the intentionally dark lenses, deviating from the clean reference.
Setting
Mask-PSNR ↑
Mixed-PSNR ↑
Mask-SSIM ↑
Mixed-SSIM ↑
Eye-ROI LPIPS ↓
Eye-NME (%) ↓
Iris-NME (%) ↓
Balanced Source Sampling
21.9821
23.3859
0.6588
0.7050
0.1125
1.4409
1.6286
w/o Eye Gradient
21.9915
23.3927
0.6592
0.7052
0.1125
1.4374
1.6241
w/o Semantic Prior
21.9883
23.3920
0.6593
0.7053
0.1122
1.4305
1.6171
w/o Multi-scale Reconstruction ( Lx )
21.3276
22.9074
0.6367
0.6939
0.1015
1.6258
1.9470
w/o Preserve Loss
21.9931
23.3945
0.6591
0.7052
0.1128
1.4360
1.6255
Ours (Natural Sampling)
22.1054
23.4710
0.6633
0.7074
0.1035
1.4030
1.5863
Appendix
Table 11 : Additional ablations of data sampling, auxiliary supervision, and semantic conditioning. All variants are evaluated on the complete paired OcuBench test set.
Setting
Mask-PSNR ↑
Mixed-PSNR ↑
Mask-SSIM ↑
Mixed-SSIM ↑
Eye-ROI LPIPS ↓
Eye-NME (%) ↓
Iris-NME (%) ↓
Out-PSNR ↑
Fixed Mixing Maps
21.8975
23.3077
0.6502
0.6967
0.0999
1.4075
1.5821
26.2977
Learned Residual Gates (Ours)
22.1054
23.4710
0.6633
0.7074
0.1035
1.4030
1.5863
26.4147
Appendix
Table 12 : Ablation of the frequency-renderer gates. Both variants use the same native-resolution frequency decomposition and synthesis; the learned variant predicts content-adaptive residual corrections to the fixed mixing maps.
Component
Time (ms)
Share (%)
Exposed CPU-preparation wait
17.6631
31.2869
Native-resolution renderer
14.1714
25.1021
pMF
12.5873
22.2962
Patch preparation
5.3085
9.4030
PIL materialization
2.8555
5.0580
Batch packing
2.4048
4.2597
Appendix
Table 14 : Steady-state runtime breakdown of OcuFlow under sustained pipelined execution ( 56.4551 ms/image).
Despite remarkable progress in reflection removal, current methods primarily exploit static image priors from a single frame and still suffer from severe residual artifacts due to the inherent ambiguity between the reflection and transmission layers. In this paper, we propose leveraging event signals to break this ambiguity. By employing event cameras to capture micro-dynamics, we reveal the differential motion between these two layers. We thereby present a novel event-driven reflection removal network, EvReflection, that utilizes these dynamic cues for layer separation. Specifically, we design a Micro-Dynamics Decoupler to disentangle layer-specific motions from event streams as priors, which then guide a Parallax-Attention Rectifier to cleanly remove artifacts from the RGB image. Furthermore, to address data scarcity, we develop a parallax-aware simulation pipeline and construct the EVR2 benchmark dataset, the first real-world dataset for this task. Extensive experiments demonstrate that EvReflection achieves state-of-the-art performance on both synthetic and real-world benchmarks, surpassing the best competing method by more than 1.6 dB and 1.2 dB in PSNR, respectively. The code, dataset, and pre-trained models are available at https://github.com/JiaxiaoWang/EvReflection.
Jiaxiao Wang, Dachun Kai, Huyue Zhu +3
University of Science and Technology of China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center
Single-image reflection removal aims to recover a clean transmission layer from one image captured through glass. We study an explicit decomposition pipeline built on RDNet and introduce LowAux, a training-only low-pass reflection auxiliary objective. The original residual target remains the main reflection supervision, while symmetrically filtered prediction and target provide a stable low-frequency constraint. We further incorporate scene-balanced real pairs from RRW to broaden real-scene coverage and improve cross-dataset generalization. To avoid evaluation discrepancies caused by model-specific resizing, padding, output quantization, and metric code, we build a unified public benchmark over CEILNet, Real20, Postcard, Objects, and Wild. Under the same evaluator, the proposed system obtains a five-dataset macro average of 27.546 dB PSNR, 0.9220 SSIM, 0.9751 NCC, and 0.004760 LMSE, achieving the highest macro-average PSNR, SSIM, and NCC and the lowest LMSE among the compared public checkpoints and internal variants. Per-dataset and qualitative analyses show that the main benefit is a more balanced performance across diverse reflection distributions, while clear semantic reflections in Postcard remain challenging.
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection_S2R.