We present ContinuityCam, a novel approach to generate a continuous video from a single static RGB image and an event camera stream. Conventional cameras struggle with high-speed motion capture due to bandwidth and dynamic range limitations. Event cameras are ideal sensors to solve this problem because they encode compressed change information at high temporal resolution. In this work, we tackle the problem of event-based continuous color video decompression, pairing single static color frames and event data to reconstruct temporally continuous videos. Our approach combines continuous long-range motion modeling with a neural synthesis model, enabling frame prediction at arbitrary times within the events. Our method only requires an initial image, thus increasing the robustness to sudden motions, light changes, minimizing the prediction latency, and decreasing bandwidth usage. We also introduce a novel single-lens beamsplitter setup that acquires aligned images and events, and a novel and challenging Event Extreme Decompression Dataset (E2D2) that tests the method in various lighting and motion profiles. We thoroughly evaluate our method by benchmarking color frame reconstruction, outperforming the baseline methods by 3.61 dB in PSNR and by 33% decrease in LPIPS, as well as showing superior results on two downstream tasks.
Figures & tables
Figure 1 : Event-based continuous color video decompression uses an initial frame and subsequent events to generate frames. The prediction relies on continuous motion estimation, neural synthesis and image generation modules.
Figure 2 : For natural camera motions, it is common to have sharp frames followed by blurry frames (shaded blue region above), which prohibits interpolation methods. Our method is able to reconstruct in these scenarios due to the removal of the dependency on the second frame.
Figure 3 : Overview . The initial frame and the long range event volumes are concatenated forming the network input. Top blue box ( Sec. 3.1 ) : A continuous motion network regresses the motion coefficients for generating the point trajectory for every pixel from events and the initial frame. Bottom orange box ( Sec. 3.2 ) : The input is projected to tri-plane features. A lightweight decoder queries the features and synthesizes pixel RGB values. Bottom red box ( Sec. 3.3 ) : Optical flow is compuated between the intial frame and the synthesized latent frame. We compute another set of features and warped images as pyramids. Right green box ( Sec. 3.4 ) : Finally, the splatted features and images are merged with the synthesized images via a mult-scale fusion network into a high-quality color image prediction.
Figure 4 : Continuous long-term trajectory output on test sequences of BS-ERGB [ 63 ] dataset. We show pixel tracks of uniformly initialized features using motion coefficients predicted from events. The network outputs dense tracks (i.e., per-pixel) in a single feedforward pass. Our continuous basis-enabled motion module can decode complex long-range motions up to 1 second.
Figure 5 : Qualitative Evaluation: We present two qualitative examples from E2D2 and BS-ERGB [ 63 ] , respectively. Our method, ContinuityCam, demonstrates enhanced accuracy in reconstructing geometry, even with challenging deformable subjects, such as in the “Fire” sequence (d). This improvement is attributed to the effective use of event data. Notably, in low-light conditions, as seen in the “Gnome” sequence (a), our approach markedly reduces motion blur compared to traditional image acquisition methods. While FILM [ 49 ] generates plausible results, it fails to accurately predict geometry in all examples. DMVFN [ 24 ] struggles with occlusions, particularly those caused by rotational movements, as evident in the “Gnome” sequence.
Figure 6 : We contribute the E2D2 dataset for benchmarking this task under challenging light conditions, using our newly designed single-lens beamsplitter.
Figure 7 : Qualitative downstream applications comparing original blurry frames (Left in each subfigure) to the decompressed frames (Right in each subfigure).
Figure 8 : Backbone architecture. The U-Net with skip connections maps event volumes and images into multi-scale dense output.
Figure 9 : The architecture for the K-planes synthesis module. The initial image and event volume are mapped to multi-scale Tri-planes (xy, xt, yt). These feature planes are sampled bilinearly and fed into a lightweight color decoder network.
Figure 10 : Architecture for the multi-scale feature fusion network. The warped feature and image pyramids are gradually injected into the network via a series of upsampling and convolution.
Figure 11 : Constructed single objective beam splitter with each major component labeled. As constructed the Event Camera will be flipped compared to a traditional setup.
School of Computer Science, Wuhan University, Wuhan, China · School of Artificial Intelligence, Ningbo University, Ningbo, China · National University of Singapore, Singapore +2
Instituto de Telecomunicações, Instituto Universitário de Lisboa (ISCTE-IUL) Lisbon, Portugal · Instituto de Telecomunicac¸˜oes, Instituto Universit´ario de Lisboa (ISCTE-IUL) Lisbon, Portugal