Learning Which Correspondences to Trust: Confidence-Weighted Event-Camera Localization in LiDAR Maps
Authors: Panagiotis Kiousis, Kuangyi Chen, Jun Zhang, Friedrich Fraundorfer
Organizations: Dept. of Electrical & Computer Engineering, Univ. of Patras, Greece. · Institute of Visual Computing, Graz University of Technology, Graz, Austria.
Localizing an event camera against a pre-built LiDAR map can be cast as dense optical-flow estimation between a rendered depth view and an event image, followed by a Perspective-n-Point (PnP) solver over the induced 3D-2D correspondences. Existing pipelines rely on geometric consensus during pose estimation, but do not explicitly model the reliability or pose informativeness, i.e., how strongly a correspondence constrains the camera pose, of individual correspondences. We show that the natural way to learn it -- using the per-correspondence error to constrain the learning of confidence -- suffers from a depth-dependent bias: small pixel errors reside predominantly at large depths and do not lead to high pose informativeness. Instead, in our method (CELL), we learn a per-correspondence confidence end-to-end through the pose, using a differentiable probabilistic PnP whose log-partition term encourages weight configurations that yield a better-constrained pose distribution. The learned confidence is used in three ways: (i) it reweights the flow supervision in a decoupled training scheme that keeps pose gradients out of the flow/edge backbone; (ii) it drives a probabilistic correspondence selection at test time; and (iii) together with the network's edge-probability it weights a final edge-matching refinement. We further design a partial-completion depth representation that adds signal without hallucinating across large gaps. On M3ED and DSEC our full system improves over the LEAR baseline on the majority of the evaluated sequences: it reduces the median translation error by up to 26.9% and the median rotation error by up to 15.8%.
Figures & tables
Fig. 1: Our pipeline. We extend the LEAR baseline with three components ( green, dashed ): a partial depth completion, a confidence head trained end-to-end through the pose loss, and a confidence-weighted edge-matching refinement . Sample maps are outputs on falcon_indoor_flight .
Fig. 2: Qualitative maps ( falcon_indoor_flight ). Top: the flow error (purple → yellow = low → high) is low almost everywhere, dominated by the depth-dependent perspective effect rather than by pose-relevant structure—a signal the confidence is not trained to reproduce. Instead, supervised only through the pose loss, the learned confidence reflects both cues: it remains high on some low-error, far-depth regions, but concentrates most strongly on near-depth, high-parallax object edges—the correspondences most informative for translation. Bottom: the edge-probability, the predicted and event edges, the pointwise product of the edge-probability and confidence maps—the refinement weight (edge-prob × conf)—and the event-edge distance transform.
Dataset
Test Scene
Time
EVLoc [ 1 ]
LEAR
Ours
T[cm] ↓
R[ ∘ ] ↓
Acc ↑
T[cm] ↓
R[ ∘ ] ↓
Acc ↑
T[cm] ↓
R[ ∘ ] ↓
Acc ↑
M3ED
car_forest_into_ponds
day
19.76
0.86
–
15.15
0.739
72.8
12.25
0.711
82.9
car_urban_day_penno
day
14.15
0.60
–
12.14
0.525
89.1
11.94
0.530
90.3
falcon_outdoor_day_penno_parking
day
25.11
1.32
–
23.61
1.186
48.8
19.84
1.000
59.2
falcon_indoor_flight
-
8.11
0.97
–
6.95
0.796
29.3
6.19
0.670
39.1
spot_indoor_building
-
12.88
2.07
–
10.85
1.718
13.6
10.53
1.672
16.7
TABLE I: Comparison with EVLoc and LEAR on M3ED and DSEC. Metrics as in Sec. V-A ; Ours : confidence-guided selection with confidence-weighted edge-matching refinement; bold: best per row. EVLoc results as reported in [ 2 ] (no accuracy).
Fig. 3: Confidence-guided selection ( falcon_indoor_flight ). Correspondences kept by each rule, coloured by confidence (green = high, red = low; subsampled for clarity), and the probability of sampling rule overlaid on the confidence histogram. floor0 keeps a broad but confidence-biased subset, whereas the hard > median cut discards half the points.
Flow-error-supervised
Pose-supervised
15.69/1.59/4.3/7.04
7.74/0.83/22.4/5.92
TABLE II: Flow-error-supervised compared to pose-supervised confidence ( falcon_indoor_flight , sparse depth): T[cm]/R[ ∘ ]/Acc/mean depth of kept points [m].
Selection
med_t
med_r
Acc
W/N ↑
Δ % ↓
W/N ↑
Δ % ↓
W/N ↑
Δ pp ↑
all
18/22
− 7.1
13/22
− 2.1
14/22
+2.4
floor0
18/22
− 7.5
15/22
− 3.2
14/22
+2.8
floor0.5
19/22
− 7.2
13/22
− 2.5
14/22
+2.5
>median
14/22
− 5.2
14/22
− 3.0
13/22
+2.4
TABLE III: Correspondence selection (non-refined poses, 22 sequences). W/N: sequences on which a rule beats the baseline; Δ : mean change (median T/R in %, Acc in points).
Fig. 4: Depth-completion variants (on falcon_indoor_flight ): sparse LiDAR ( ∼ 20%), our edge-preserving partial completion ( ∼ 68%), and full completion ( ∼ 99.8%). Full propagates depth across object boundaries (red box), unlike partial (Table V ).
Edge weighting
Mean Acc
Δ vs LEAR
Δ vs non-ref.
≥ base
none (uniform)
83.18
+3.06
+0.26
16/22
conf (confidence)
83.21
+3.09
+0.29
16/22
both (edge-prob × conf)
83.24
+3.12
+0.32
16/22
TABLE IV: Edge weighting for the refinement : mean accuracy and its gain over LEAR and the non-refined pose.
Depth input
Valid
med_t [cm] ↓
acc O [%] ↑
EPE mn↓
EPE md↓
sparse
∼ 20%
23.15
51.1
8.74
6.06
full dense
∼ 99.8%
23.26
50.6
10.23
6.42
partial (ours)
∼ 68%
20.80
57.6
8.08
5.78
TABLE V: Depth-input completion on falcon_outdoor_day_penno_parking (all correspondences, no selection).
Configuration
med_t ↓
med_r ↓
Acc ↑
EPE mn↓
EPE md↓
base: LEAR (sparse)
23.61
1.186
48.8
9.42
5.97
+ partial depth
22.65 ( − 0.96)
1.240 (+0.054)
51.1 (+2.3)
9.63
6.31
+ decoupled conf. training
20.80 ( − 1.85)
1.080 ( − 0.160)
57.6 (+6.5)
8.08
5.78
+ confidence selection
20.06 ( − 0.74)
1.027 ( − 0.053)
58.9 (+1.3)
7.48
5.27
+ edge-matching refinement ( Ours )
19.84 ( − 0.22)
1.000 ( − 0.027)
59.2 (+0.3)
= floor0 †
net (base → Ours)
19.84 ( − 3.77)
1.000 ( − 0.186)
59.2 (+10.4)
7.48
5.27
TABLE VI: Incremental ablation on falcon_outdoor_day_penno_parking : each row adds one component; parentheses give the change vs. the previous row. † Refinement changes only the pose, so EPE equals the floor0 row.
Split
med_t
med_r
Acc
W/N ↑
Δ % ↓
W/N ↑
Δ % ↓
W/N ↑
Δ pp ↑
M3ED (9)
6/9
− 3.2
6/9
− 1.8
7/9
+3.3
DSEC (13)
7/13
− 3.2
7/13
− 1.6
5/13
− 0.2
All (22)
13/22
− 3.2
13/22
− 1.7
12/22
+1.2
TABLE VII: Halved iteration budget : ours at 12 IFR iterations vs. the 24-iteration LEAR baseline (W/N and Δ as in Table III ).
Method
Stage
Runtime [ms]
Memory [MB]
Baseline (LEAR)
Feature encoding
8.89
335
Flow estimation (24 iters)
77.15
Ours
Feature encoding
8.90
339
Flow estimation (24/12 iters)
76.24/39.15
Confidence head
0.09
Edge-matching refinement
203.82
TABLE VIII: Runtime and memory on falcon_indoor_flight ; memory is the peak allocation, excluding the fixed CUDA context ( ≈ 300 MB).
Event cameras capture brightness changes asynchronously with microsecond resolution, yet existing optical flow methods fail to fully exploit this temporal continuity. Frame-based approaches impose artificial accumulation latency and suffer from domain overfitting, while model-based local methods operate statelessly, discarding temporal history between predictions and yielding inaccurate flows. We propose \textbf{LC-Flow}, the first temporally continuous, learning-based optical flow estimator that operates purely from local events. At its core, a Continuous Local Recurrent Network maintains persistent hidden states per spatial grid, incrementally accumulating temporal context as events arrive. Unlike frame-based methods constrained to fixed accumulation windows, and unlike stateless model-based methods that recompute motion from scratch at each step, LC-Flow produces sparse local flow estimates at arbitrary timestamps with full motion history. To address the inherent ambiguity of local observations, we jointly learn a confidence score that quantifies the reliability of each prediction, explicitly handling event sparsity and the aperture problem. This confidence serves a dual role: filtering unreliable estimates for downstream tasks such as visual odometry, and providing principled weights for a multi-scale confidence-guided aggregation that reconstructs globally consistent flow from the sparse local outputs. LC-Flow achieves state-of-the-art performance among local methods on both MVSEC and DSEC, while the confidence-guided aggregation establishes a new overall state-of-the-art on the MVSEC benchmark, surpassing heavy frame-based networks that rely on global spatial priors.
Image-to-Point Cloud Registration (I2P) is essential for integrating camera and LiDAR in perception and autonomous systems, yet the modality gap between images and point clouds makes it difficult to achieve both high accuracy and strong generalization. In this paper, we propose a simple yet effective I2P method that treats LiDAR as an imaging sensor: from a single sparse LiDAR scan, we generate a dense LiDAR intensity image using Conditional Rectified Flow, match it with a camera image using a pre-trained feature matcher, and estimate the 6-DoF relative pose via PnP-RANSAC. The proposed model is pre-trained through a self-supervised image completion task and fine-tuned on a small amount of LiDAR data (neither image-point cloud pairs nor ground-truth sensor poses are required), enabling it to scale to diverse LiDAR and camera configurations. Experiments on the R3LIVE dataset show that the proposed method achieves a mean error of 4.89° / 1.63 m, outperforming existing methods, while completing a single registration in approximately 0.68 s.
Reon Tabata, Kenji Koide, Shuji Oishi +4
Department of Computer Science and Engineering, Toyohashi University of Technology, Toyohashi, Aichi, Japan · National Institute of Advanced Industrial Science and Technology, Tsukuba, Ibaraki, Japan
LiDAR-based Scene Coordinate Regression (SCR) maps point clouds directly to 3D scene coordinates, enabling precise 6-DoF localisation without explicit map retrieval. However, existing methods produce deterministic predictions, discarding aleatoric uncertainty that could improve robustness and downstream decision-making. We present UQ-Loc, which extends the LightLoc architecture with an anisotropic Gaussian covariance head that predicts a full 3x3 positive-definite covariance matrix per voxel. Training uses a Negative Log-Likelihood (NLL) loss augmented with a kNN-based spatial smoothness regulariser, while inference employs a modified SC2-PCR solver with uncertainty-weighted seed scoring and a Mahalanobis-distance inlier test. We adopt Expected Calibration Error (ECE) as a principled metric for evaluating the quality of the predicted uncertainty. Experiments demonstrate that UQ-Loc achieves consistent improvement in 6-DoF localization accuracy while producing well-calibrated covariances.