Temporally dense optical flow is essential for dynamic perception in immersive VR/AR systems, where rapid head, hand, and object motion must be continuously captured and tracked. Existing frame-based optical flow estimation methods are constrained by the tradeoff between temporal resolution and computational cost; while event cameras, with their high temporal resolution and energy efficiency, serve as a natural solution to the dilemma. However, event-based approaches commonly rely on correlation volumes to capture pairwise voxel correspondences, which incur substantial memory and computation overhead. We present E-WAVE, a correlation-free framework for high-temporal-resolution (HTR) optical flow estimation from event streams. Instead of constructing all-pairs correlation volumes, E-WAVE employs global attention mechanism to model long-range feature dependencies and performs trajectory guided feature warping using Bézier curve. Through iterative updates, it predicts trajectories that allow for querying at arbitrary timestamps without repeated inference. Experiments on MultiFlow and DSEC-Flow demonstrate a 25% lower trajectory error and comparable endpoint flow estimation accuracy relative to state-of-the art baselines. Additional evaluations on self-captured data using a head-mounted prototype validate that E-WAVE remains robust under challenging real-world conditions.
Figures & tables
Figure 1: Correlation-based vs. Warping-based matching. (a) Correlation volumes store dense pixel-wise similarity between two feature frames, which is queried at a later stage to update estimated trajectories. (b) Our method leverages global attention to capture long-range correlation and directly derive warping parameters.
Figure 2: Pipeline of \method . We firstly extract features from voxelized events and RGB frames (if available), and then repeat the refinement step by R times. At each iteration, visual features are warped using the predicted trajectory map from the previous iteration (purple), concatenated with the trajectory map and the hidden state (cyan), and processed by a ViT-DPT module to obtain an updated hidden state. Ultimately, two convolution heads convert the hidden state into the updated trajectory map and uncertainty vector ρ , respectively. An NLL loss on endpoints and a trajectory loss on intermediate ground truth are computed for each iteration to train the entire framework.
Figure 3: Visualization of estimated optical flow and predicted (red) vs. ground truth (blue) trajectories across different iterations. The warp-and-update process improves both pixelwise trajectory alignment and object-level flow coherence in the first few iterations and converges afterwards.
Figure 4: Qualitative comparison on DSEC-Flow [ 15 ] . \method produces more detailed estimation thanks to its higher spatial resolution. We cropped bottom 60px of the figure for the lack of supervision in this region.
Method
1PE
2PE
3PE
EPE
AE
HTR
E-RAFT
12.74
4.74
2.68
0.79
2.85
TMA
10.86
3.97
2.30
0.74
2.68
IDNet
10.07
3.50
2.04
0.72
2.72
ECDDP
8.89
3.20
1.96
0.70
2.58
BAT
7.54
2.84
1.74
0.65
2.43
ResFlow
11.22
4.24
2.50
0.75
2.73
✓
Table 1: Quantitative results on DSEC-Flow. Methods up to “Ours” are event-only, while STFlow, BFlow and “Ours w/ img” works on event-RGB inputs. “+ bidir.” and “+ finetuning” correspond to two extensions in Sec. 2.6 , respectively.
Methods
Input
TEPE
TAE
EPE
AE
HTR
WAFT
i
6.20
17.85
3.85
3.36
IDNet
e
7.08
20.74
7.60
10.85
E-RAFT
e
6.21
18.74
4.35
5.7
TMA
e
6.02
18.24
3.91
4.84
BFlow
e
2.74
7.00
5.00
7.35
✓
ResFlow
e
2.28
5.32
3.74
4.70
✓
Table 2: Quantitative results on MultiFlow. STFlow † denotes STFlow model trained with Bezier curve motion representation.
Figure 5: Qualitative results on MultiFlow. For each scene, the first row visualizes the RGB reference image, the ground truth optical flow, and estimated flows from three methods, overlaid with ground truth (blue) and estimated (red) trajectories and zoomed in to a representative region. The second row visualizes the integrated event frame, the ground truth IWE, and IWEs for each method, also zoomed in to the same region.
Methods
Trajectory
Endpoint
Ltraj
Warp
TEPE
TAE
EPE
AE
✓
B
1.81
3.89
3.03
3.51
×
B
7.62
24.60
4.87
5.34
✓
L
4.64
14.64
6.84
7.07
×
L
6.27
18.32
4.70
5.14
✓
no
3.21
6.96
5.58
6.51
Table 3: Ablation study results. “B” denotes warping with degree-10 Bézier curves, “L” denotes warping with linear trajectories, “no” denotes prediction without warping process
Resolution
Trajectory
Endpoint
Spatial
Temporal
TEPE
TAE
EPE
AE
1/2
20
1.81
3.89
3.03
3.51
1/4
20
2.02
4.73
3.39
4.28
1/8
20
2.46
6.09
3.91
4.79
1/2
20
1.81
3.89
3.03
3.51
1/2
10
2.05
4.46
3.39
3.98
Table 4: Ablation study results on spatial and temporal resolutions. Spatial resolutions are written with respect to the original data resolution, and temporal resolution is denoted by the number of event voxels used for prediction.
Figure 6: Results on real-world acquisitions. From left to right: reference RGB frames (not used for inference), event voxels at the beginning and end of inference window, and estimated results by three methods. For scenarios 2&4, we additionally visualize pixel trajectories (blue) to intuitively compare motion estimation accuracy across models.