Temporally dense optical flow is essential for dynamic perception in immersive VR/AR systems, where rapid head, hand, and object motion must be continuously captured and tracked. Existing frame-based optical flow estimation methods are constrained by the tradeoff between temporal resolution and computational cost; while event cameras, with their high temporal resolution and energy efficiency, serve as a natural solution to the dilemma. However, event-based approaches commonly rely on correlation volumes to capture pairwise voxel correspondences, which incur substantial memory and computation overhead. We present E-WAVE, a correlation-free framework for high-temporal-resolution (HTR) optical flow estimation from event streams. Instead of constructing all-pairs correlation volumes, E-WAVE employs global attention mechanism to model long-range feature dependencies and performs trajectory guided feature warping using Bézier curve. Through iterative updates, it predicts trajectories that allow for querying at arbitrary timestamps without repeated inference. Experiments on MultiFlow and DSEC-Flow demonstrate a 25% lower trajectory error and comparable endpoint flow estimation accuracy relative to state-of-the art baselines. Additional evaluations on self-captured data using a head-mounted prototype validate that E-WAVE remains robust under challenging real-world conditions.
Figures & tables
Figure 1: Correlation-based vs. Warping-based matching. (a) Correlation volumes store dense pixel-wise similarity between two feature frames, which is queried at a later stage to update estimated trajectories. (b) Our method leverages global attention to capture long-range correlation and directly derive warping parameters.
Figure 2: Pipeline of \method . We firstly extract features from voxelized events and RGB frames (if available), and then repeat the refinement step by R times. At each iteration, visual features are warped using the predicted trajectory map from the previous iteration (purple), concatenated with the trajectory map and the hidden state (cyan), and processed by a ViT-DPT module to obtain an updated hidden state. Ultimately, two convolution heads convert the hidden state into the updated trajectory map and uncertainty vector ρ , respectively. An NLL loss on endpoints and a trajectory loss on intermediate ground truth are computed for each iteration to train the entire framework.
Figure 3: Visualization of estimated optical flow and predicted (red) vs. ground truth (blue) trajectories across different iterations. The warp-and-update process improves both pixelwise trajectory alignment and object-level flow coherence in the first few iterations and converges afterwards.
Figure 4: Qualitative comparison on DSEC-Flow [ 15 ] . \method produces more detailed estimation thanks to its higher spatial resolution. We cropped bottom 60px of the figure for the lack of supervision in this region.
Method
1PE
2PE
3PE
EPE
AE
HTR
E-RAFT
12.74
4.74
2.68
0.79
2.85
TMA
10.86
3.97
2.30
0.74
2.68
IDNet
10.07
3.50
2.04
0.72
2.72
ECDDP
8.89
3.20
1.96
0.70
2.58
BAT
7.54
2.84
1.74
0.65
2.43
ResFlow
11.22
4.24
2.50
0.75
2.73
✓
Table 1: Quantitative results on DSEC-Flow. Methods up to “Ours” are event-only, while STFlow, BFlow and “Ours w/ img” works on event-RGB inputs. “+ bidir.” and “+ finetuning” correspond to two extensions in Sec. 2.6 , respectively.
Methods
Input
TEPE
TAE
EPE
AE
HTR
WAFT
i
6.20
17.85
3.85
3.36
IDNet
e
7.08
20.74
7.60
10.85
E-RAFT
e
6.21
18.74
4.35
5.7
TMA
e
6.02
18.24
3.91
4.84
BFlow
e
2.74
7.00
5.00
7.35
✓
ResFlow
e
2.28
5.32
3.74
4.70
✓
Table 2: Quantitative results on MultiFlow. STFlow † denotes STFlow model trained with Bezier curve motion representation.
Figure 5: Qualitative results on MultiFlow. For each scene, the first row visualizes the RGB reference image, the ground truth optical flow, and estimated flows from three methods, overlaid with ground truth (blue) and estimated (red) trajectories and zoomed in to a representative region. The second row visualizes the integrated event frame, the ground truth IWE, and IWEs for each method, also zoomed in to the same region.
Methods
Trajectory
Endpoint
Ltraj
Warp
TEPE
TAE
EPE
AE
✓
B
1.81
3.89
3.03
3.51
×
B
7.62
24.60
4.87
5.34
✓
L
4.64
14.64
6.84
7.07
×
L
6.27
18.32
4.70
5.14
✓
no
3.21
6.96
5.58
6.51
Table 3: Ablation study results. “B” denotes warping with degree-10 Bézier curves, “L” denotes warping with linear trajectories, “no” denotes prediction without warping process
Resolution
Trajectory
Endpoint
Spatial
Temporal
TEPE
TAE
EPE
AE
1/2
20
1.81
3.89
3.03
3.51
1/4
20
2.02
4.73
3.39
4.28
1/8
20
2.46
6.09
3.91
4.79
1/2
20
1.81
3.89
3.03
3.51
1/2
10
2.05
4.46
3.39
3.98
Table 4: Ablation study results on spatial and temporal resolutions. Spatial resolutions are written with respect to the original data resolution, and temporal resolution is denoted by the number of event voxels used for prediction.
Figure 6: Results on real-world acquisitions. From left to right: reference RGB frames (not used for inference), event voxels at the beginning and end of inference window, and estimated results by three methods. For scenarios 2&4, we additionally visualize pixel trajectories (blue) to intuitively compare motion estimation accuracy across models.
Estimating continuous optical flow is a fundamental yet challenging problem in dynamic visual perception. Event-based cameras, with microsecond latency and high dynamic range, capture brightness changes asynchronously, offering a unique opportunity to model motion with fine temporal precision. However, the scarcity of temporally dense ground-truth annotations limits the effectiveness of supervised learning, while contrast maximization (CM) frameworks, focused on sharpening the Image of Warped Events (IWE), often neglect temporal continuity and structural coherence, leading to distorted trajectories under complex motion. To overcome these challenges, we propose a hybrid-supervised framework for continuous-time optical flow estimation, grounded in the principle of Spatio-temporal Structural Consistency (STSC). This paradigm jointly enforces local structural stability and trajectory continuity, ensuring physically coherent motion across time. To further enhance representation and robustness, we design a bidirectionally complementary multi-scale architecture and employ a curriculum-guided hybrid training strategy, enabling a smooth transition from supervised point constraints to self-supervised manifold regularization. Comprehensive experiments across multiple benchmarks show that our method achieves state-of-the-art performance in both continuous-time and standard optical flow estimation, demonstrating the effectiveness of the proposed learning paradigm.
Event cameras capture brightness changes asynchronously with microsecond resolution, yet existing optical flow methods fail to fully exploit this temporal continuity. Frame-based approaches impose artificial accumulation latency and suffer from domain overfitting, while model-based local methods operate statelessly, discarding temporal history between predictions and yielding inaccurate flows. We propose \textbf{LC-Flow}, the first temporally continuous, learning-based optical flow estimator that operates purely from local events. At its core, a Continuous Local Recurrent Network maintains persistent hidden states per spatial grid, incrementally accumulating temporal context as events arrive. Unlike frame-based methods constrained to fixed accumulation windows, and unlike stateless model-based methods that recompute motion from scratch at each step, LC-Flow produces sparse local flow estimates at arbitrary timestamps with full motion history. To address the inherent ambiguity of local observations, we jointly learn a confidence score that quantifies the reliability of each prediction, explicitly handling event sparsity and the aperture problem. This confidence serves a dual role: filtering unreliable estimates for downstream tasks such as visual odometry, and providing principled weights for a multi-scale confidence-guided aggregation that reconstructs globally consistent flow from the sparse local outputs. LC-Flow achieves state-of-the-art performance among local methods on both MVSEC and DSEC, while the confidence-guided aggregation establishes a new overall state-of-the-art on the MVSEC benchmark, surpassing heavy frame-based networks that rely on global spatial priors.
Event cameras capture intensity changes asynchronously with high temporal resolution, requiring novel preprocessing methods for downstream tasks. Unlike static intensity snapshots, event data inherently encode information about scene dynamics and object motion, meaning that features derived from events can exhibit behaviors with no direct analogue in frame-based vision. In this paper, we analyze two features used in event-based corner detection---the eigenvalues of the structure tensor and the spatiotemporal density values---and show that they are \emph{motion cues}. We hypothesize that these features, combined with local geometric information, can enhance motion estimation tasks. To validate this, we first theoretically analyze how the eigenvalues of the structure tensor at moving corner points relate to the direction of motion. We then design controlled experiments on a synthetic dataset, confirming that extending local geometric features with eigenvalues and density values provides complementary motion information and is robust to texture and shot noise. Finally, we integrate the proposed features into a state-of-the-art event-based optical flow network and evaluate on the real-world DSEC benchmark, where the added features consistently improve accuracy, with the largest gains in data-scarce scenarios and for lower-capacity models. The code for this paper can be found at: \href{https://github.com/hesamaraghi/static-in-frames-dynamic-in-events}{https://github.com/hesamaraghi/static-in-frames-dynamic-in-events}.
Hesam Araghi, Jan van Gemert, Nergis Tomen
Computer Vision Lab, Delft University of Technology