Organizations: Hong Kong University of Science and Technology (Guangzhou) · National University of Singapore · Xidian University · Tsinghua University · Huazhong University of Science and Technology · Peking University · The University of Tokyo · Robotics and Perception Group, University of Zurich
Event cameras provide a high dynamic range and preserve brightness-change cues in lighting conditions where conventional RGB frames may be noisy or saturated. To benchmark event-guided restoration across a broad illumination range, we organized the SEE Challenge 2026 with the Event-Based Multimodal Vision Workshop at ECCV 2026. The task conditions restoration on one or more RGB frames, synchronized events, and a scalar target-brightness statistic provided by the organizers. It uses SEE-600K, which contains 610,126 image-event observations from 202 real-world scenes spanning low-light, normal-light, and high-light conditions with illumination variations of up to 1,000×. The challenge follows an open-system protocol: participants may use different temporal contexts, architectures, pretrained weights, test-time augmentation, and post-processing strategies. PSNR determines the ranking, and SSIM is reported as a secondary metric. Around 70 teams registered interest and 15 valid CodaBench submissions were received. Six distinct teams completed organizer-side identity and technical verification, provided method descriptions, checkpoints, inference code, and instructions, and are included in the verified open-system ranking reported here. Beyond the ranking, this report analyzes exposure subsets, semantically distinct test cases, a shared failure pattern, system design choices, and inference strategies. The top systems obtain closely spaced average scores, while the best-performing method varies across cases and metrics; under severe underexposure, all verified systems retain visible local errors.
Figures & tables
Figure 1 : Adapted from SEE-Net [ 15 ] . Brightness distributions and task settings of SDE [ 4 ] and SEE-600K [ 15 ] . (a) SDE covers a low-to-normal brightness range. (b) SEE-600K spans low-light, normal-light, and high-light conditions. (c) Earlier low-light methods map dark inputs to a fixed normal-light target. (d) The official SEE Challenge evaluation conditions restoration on a target-brightness statistic B derived from the reference image.
Split
Scenes
Pairs
Reference
Purpose
Motion setting
Training split
165
509,132
Public
Training
Camera motion
Test split
37
147,368
Public
Phase 1 validation
Camera motion
Hidden final split
6
3,405
Hidden
Phase 2 validation
Camera and scene motion
Table 1 : Data used by the challenge. Pair counts refer to constructed source–target restoration pairs rather than distinct raw frames. The three splits are scene-disjoint.
(a) Phase 2 evaluation on Codabench
Participant
Average
Low-to-Normal
High-to-Normal
FLOPs (G)
Params (M)
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
pixelartai
23.1810
0.7246
21.2066
0.6965
25.1555
0.7527
1314.77
10.24
vincent2013
23.0865
0.7178
21.1040
0.6941
25.0689
0.7416
3390.56
54.21
EventFuse
22.7992
0.7057
21.1942
0.6809
24.4042
0.7305
97.51
1.16
bucloud
22.6168
0.7048
20.7068
0.6779
24.5268
0.7316
870.37
15.85
Table 2 : Results of the SEE Challenge 2026 under Phase 2 Codabench and organizer-side dense sampling. Dense-sampling PSNR determines the ranking, while SSIM is secondary. Best and second-best results are shown in bold and underlined . FLOPs (G) and parameters are reported at the native resolution. FLOPs (G) include cascaded stages but exclude TTA, temporal averaging, ensembles, and post-processing.
Robot Motion (G1)
Checkerboard (G2)
Chinese Text (G3)
Team
PSNR ↑
SSIM ↑
L1 ↓
LPIPS ↓
PSNR ↑
SSIM ↑
L1 ↓
LPIPS ↓
PSNR ↑
SSIM ↑
L1 ↓
LPIPS ↓
pixelartai
18.76
0.4998
0.0826
0.2767
21.49
0.7951
0.0638
0.1553
24.83
0.7486
0.0375
0.1838
vincent2013
18.85
0.5010
0.0805
0.2701
21.32
0.7941
0.0657
0.1548
25.06
0.7583
0.0364
0.1800
EventFuse
18.73
0.4889
0.0847
0.3100
21.80
0.7902
0.0622
0.1556
24.40
0.7498
0.0427
0.1877
bucloud
18.71
0.4956
0.0848
0.3069
21.17
0.7814
0.0668
0.1694
24.52
0.7484
0.0400
0.1777
jiangbin
18.79
0.4986
0.0852
0.3431
21.11
0.7603
0.0706
0.1748
23.69
0.7299
0.0467
0.2217
Table 3 : Performance comparison across different groups.
Figure 2 : Qualitative comparison of the verified submissions on representative Low-to-Normal and High-to-Normal examples. Input RGB frames, event visualizations, reference images, restored outputs, and enlarged difficult regions are displayed with a common image range and without additional gamma or contrast adjustment. The examples illustrate differences in residual noise, color, local contrast, and texture recovery that are not fully captured by average PSNR and SSIM.
Figure 3 : Representative shared failure case under severe underexposure. (a) Accumulated events, (b) input frame, (c) SEE-Net baseline, (d)–(e) and (g)–(j) verified submissions, and (f) ground-truth frame. The green boxes indicate the enlarged regions shown below. Although the submitted systems substantially brighten the input, none faithfully reconstructs the checkerboard structure. The values below each result denote PSNR/SSIM/LPIPS ( ↑/↑/↓ ).
Team
Pretraining
Prompt
TTA
Temporal Processing
Effective Passes
pixelartai
×
✓
×
✓
1
vincent2013
✓
✓
✓
×
16
EventFuse
×
✓
×
×
1
bucloud
×
✓
✓
×
4
jiangbin
✓
✓
×
×
1
aka18
×
✓
✓
×
4
Table 4 : Comparison of training and inference strategies adopted by the participating teams and the SEE baseline. All methods use the brightness prompt. Effective Passes denotes the number of model forward passes or iterative updates used during inference.
Event-based cameras are bio-inspired sensors that detect light changes asynchronously for each pixel. They are increasingly used in fields like computer vision and robotics because of several advantages over traditional frame-based cameras, such as high temporal resolution, low latency, and high dynamic range. As with any camera, the output's quality depends on how well the camera's settings, called biases for event-based cameras, are configured. While frame-based cameras have advanced automatic configuration algorithms, there are very few such tools for tuning these biases. A systematic testing framework would require observing the same scene with different biases, which is tricky since event cameras only generate events when there is movement. Event simulators exist, but since biases heavily depend on the electrical circuit and the pixel design, available simulators are not well suited for bias tuning. To allow reproducibility, we present BiasBench, a novel event dataset containing multiple scenes with settings sampled in a grid-like pattern. We present three different scenes, each with a quality metric of the downstream application. Additionally, we present a novel, RL-based method to facilitate online bias adjustments.
Andreas Ziegler, David Joseph, Thomas Gossard +2
Cognitive Systems Group, University of Tübingen, Germany
Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking the essential global illumination in images and the inherent noise sensitivity of event signals in real-world scenarios. To address these issues, we propose EIC-LIE, an event-illumination collaborative LIE framework. Concretely, we first design an Event-Illumination Collaborative Interaction (EICI) module, which contains two key processes: forward gathering, which gathers HDR features across varying lighting conditions, and backward injection, which provides complementary content for illumination and event representations. Next, we introduce an Illumination-aware Event Filter (IAEF) that dynamically reduces event noise based on brightness statistics derived from images. Additionally, we build a beam-splitter-based hybrid imaging system to collect high-quality event-image pairs with temporal synchronization from dynamic scenes, providing the first high-resolution, real-world event-based LIE dataset. Extensive experiments show that our EIC-LIE outperforms state-of-the-art methods on five real-world and synthetic datasets, significantly surpassing previous methods with improvements of up to 1.24dB in PSNR and 0.069 in SSIM. The code and dataset are released at https://github.com/QUEAHREN/EIC-LIE.
Low-light image enhancement is severely ill-posed when the input frame contains missing structure, saturated noise, and weak local contrast. Event cameras provide asynchronous brightness-change observations with high temporal resolution, but prior works often treat voxel channels as an unordered or static feature stack before fusion, rather than explicitly modeling their within-window temporal evolution, weakening the temporal evidence that makes events useful. We propose EvLIR, a temporal-residual enhancement framework that learns illumination residuals from ordered events for low-light image enhancement. Given a low-light frame and its aligned event voxel, EvLIR preserves the ordered temporal bins of the event stream and introduces a Temporal Event Residual Module (TERM) to encode short-window event dynamics with a lightweight ConvGRU. The resulting temporal state is converted into a bounded illumination correction, which provides spatially adaptive photometric guidance for Retinex-style illumination estimation and subsequent reliability-aware image-event restoration. On SDE and SDSD indoor/outdoor benchmarks, EvLIR achieves the best result on eleven of twelve dataset-metric pairs, with average scores of 25.63dB PSNR, 28.30dB PSNR*, and 0.827 SSIM across the four benchmarks.
Haoxian Zhou, Chuanzhi Xu, Langyi Chen +5
1The University of Sydney · 2Massachusetts Institute of Technology