Organizations: Hong Kong University of Science and Technology (Guangzhou) · National University of Singapore · Xidian University · Tsinghua University · Huazhong University of Science and Technology · Peking University · The University of Tokyo · Robotics and Perception Group, University of Zurich
Event cameras provide a high dynamic range and preserve brightness-change cues in lighting conditions where conventional RGB frames may be noisy or saturated. To benchmark event-guided restoration across a broad illumination range, we organized the SEE Challenge 2026 with the Event-Based Multimodal Vision Workshop at ECCV 2026. The task conditions restoration on one or more RGB frames, synchronized events, and a scalar target-brightness statistic provided by the organizers. It uses SEE-600K, which contains 610,126 image-event observations from 202 real-world scenes spanning low-light, normal-light, and high-light conditions with illumination variations of up to 1,000×. The challenge follows an open-system protocol: participants may use different temporal contexts, architectures, pretrained weights, test-time augmentation, and post-processing strategies. PSNR determines the ranking, and SSIM is reported as a secondary metric. Around 70 teams registered interest and 15 valid CodaBench submissions were received. Six distinct teams completed organizer-side identity and technical verification, provided method descriptions, checkpoints, inference code, and instructions, and are included in the verified open-system ranking reported here. Beyond the ranking, this report analyzes exposure subsets, semantically distinct test cases, a shared failure pattern, system design choices, and inference strategies. The top systems obtain closely spaced average scores, while the best-performing method varies across cases and metrics; under severe underexposure, all verified systems retain visible local errors.
Figures & tables
Figure 1 : Adapted from SEE-Net [ 15 ] . Brightness distributions and task settings of SDE [ 4 ] and SEE-600K [ 15 ] . (a) SDE covers a low-to-normal brightness range. (b) SEE-600K spans low-light, normal-light, and high-light conditions. (c) Earlier low-light methods map dark inputs to a fixed normal-light target. (d) The official SEE Challenge evaluation conditions restoration on a target-brightness statistic B derived from the reference image.
Split
Scenes
Pairs
Reference
Purpose
Motion setting
Training split
165
509,132
Public
Training
Camera motion
Test split
37
147,368
Public
Phase 1 validation
Camera motion
Hidden final split
6
3,405
Hidden
Phase 2 validation
Camera and scene motion
Table 1 : Data used by the challenge. Pair counts refer to constructed source–target restoration pairs rather than distinct raw frames. The three splits are scene-disjoint.
(a) Phase 2 evaluation on Codabench
Participant
Average
Low-to-Normal
High-to-Normal
FLOPs (G)
Params (M)
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
pixelartai
23.1810
0.7246
21.2066
0.6965
25.1555
0.7527
1314.77
10.24
vincent2013
23.0865
0.7178
21.1040
0.6941
25.0689
0.7416
3390.56
54.21
EventFuse
22.7992
0.7057
21.1942
0.6809
24.4042
0.7305
97.51
1.16
bucloud
22.6168
0.7048
20.7068
0.6779
24.5268
0.7316
870.37
15.85
Table 2 : Results of the SEE Challenge 2026 under Phase 2 Codabench and organizer-side dense sampling. Dense-sampling PSNR determines the ranking, while SSIM is secondary. Best and second-best results are shown in bold and underlined . FLOPs (G) and parameters are reported at the native resolution. FLOPs (G) include cascaded stages but exclude TTA, temporal averaging, ensembles, and post-processing.
Robot Motion (G1)
Checkerboard (G2)
Chinese Text (G3)
Team
PSNR ↑
SSIM ↑
L1 ↓
LPIPS ↓
PSNR ↑
SSIM ↑
L1 ↓
LPIPS ↓
PSNR ↑
SSIM ↑
L1 ↓
LPIPS ↓
pixelartai
18.76
0.4998
0.0826
0.2767
21.49
0.7951
0.0638
0.1553
24.83
0.7486
0.0375
0.1838
vincent2013
18.85
0.5010
0.0805
0.2701
21.32
0.7941
0.0657
0.1548
25.06
0.7583
0.0364
0.1800
EventFuse
18.73
0.4889
0.0847
0.3100
21.80
0.7902
0.0622
0.1556
24.40
0.7498
0.0427
0.1877
bucloud
18.71
0.4956
0.0848
0.3069
21.17
0.7814
0.0668
0.1694
24.52
0.7484
0.0400
0.1777
jiangbin
18.79
0.4986
0.0852
0.3431
21.11
0.7603
0.0706
0.1748
23.69
0.7299
0.0467
0.2217
Table 3 : Performance comparison across different groups.
Figure 2 : Qualitative comparison of the verified submissions on representative Low-to-Normal and High-to-Normal examples. Input RGB frames, event visualizations, reference images, restored outputs, and enlarged difficult regions are displayed with a common image range and without additional gamma or contrast adjustment. The examples illustrate differences in residual noise, color, local contrast, and texture recovery that are not fully captured by average PSNR and SSIM.
Figure 3 : Representative shared failure case under severe underexposure. (a) Accumulated events, (b) input frame, (c) SEE-Net baseline, (d)–(e) and (g)–(j) verified submissions, and (f) ground-truth frame. The green boxes indicate the enlarged regions shown below. Although the submitted systems substantially brighten the input, none faithfully reconstructs the checkerboard structure. The values below each result denote PSNR/SSIM/LPIPS ( ↑/↑/↓ ).
Team
Pretraining
Prompt
TTA
Temporal Processing
Effective Passes
pixelartai
×
✓
×
✓
1
vincent2013
✓
✓
✓
×
16
EventFuse
×
✓
×
×
1
bucloud
×
✓
✓
×
4
jiangbin
✓
✓
×
×
1
aka18
×
✓
✓
×
4
Table 4 : Comparison of training and inference strategies adopted by the participating teams and the SEE baseline. All methods use the brightness prompt. Effective Passes denotes the number of model forward passes or iterative updates used during inference.