Authors: Daehan Lee, Hyungtae Lim, Seongjun Kim, Soonbin Rho, Changhyeon Lee, Sanghyun Park, Junwoo Hong, Eunseon Choi, +2 more
Organizations: Computational Control Engineering Laboratory (CoCEL), Department of Convergence IT Engineering and Electrical Engineering, Pohang University of Science and Technology (POSTECH), Pohang 37673, South Korea · Laboratory for Information & Decision Systems (LIDS), Massachusetts Institute of Technology, Cambridge, MA 02139, USA
For field robotic missions such as inspection, search-and-rescue, and exploration, light detection and ranging (LiDAR)-inertial odometry (LIO) can serve as a core component of autonomy by providing localization and mapping in GNSS-denied or unstructured environments. However, transitions between confined and open spaces, which are commonly encountered in field deployments, can induce substantial changes in scan density and local geometric structure, thereby reducing the robustness and computational efficiency of LIO. To address these issues, we present GenZ-LIO, a generalizable LIO framework designed to adapt to variations in spatial scale across confined and open environments. GenZ-LIO comprises three components: (i) scale-aware adaptive voxelization for regulating scan downsampling across spatial-scale changes, (ii) hybrid-metric state update for combining point-to-plane and point-to-point residuals under varying geometric structure, and (iii) voxel-pruned correspondence search for efficient point-to-point matching. We conduct a comprehensive evaluation using 42 sequences from nine public datasets and our newly collected NarrowWide dataset to analyze LIO performance under spatial-scale variations across diverse field scenarios. Across the evaluated sequences, GenZ-LIO maintains stable odometry estimation without divergence, indicating practical robustness under the tested field conditions. Our code is available at https://github.com/cocel-postech/genz-lio .
Figures & tables
Figure 1 : Estimated trajectory and mapping result of our GenZ-LIO on the Handheld-A-01 sequence of our NarrowWide dataset, introduced in this work, where the platform traverses environments with substantially different spatial scales. The trajectory is color-coded by the proposed scale indicator mˉt , which reflects the spatial extent of the surrounding scene. By incorporating this indicator, GenZ-LIO adapts to varying spatial scales, enabling consistent odometry estimation while maintaining computational efficiency across both confined and open areas.
Figure 2 : Comparison of adaptive voxelization strategies. (a) LOCUS 2.0 [ 34 ] updates the voxel size for the next scan, denoted by dt+1 , using the ratio of the current voxelized point count Nt to the fixed desired point count Ndesiredfixed . (b) AdaLIO [ 26 ] adaptively selects between the predefined coarse voxel size dcoarsefixed and the fine voxel size dfinefixed by checking whether the coarse voxelization yields fewer points than a fixed threshold τNfixed . (c) LIVOX-CAM [ 9 ] adjusts the voxel size dt by first performing a temporary voxelization with a fixed voxel size dtempfixed and then updating it based on the ratio between the temporary point count Ntemp,t and the desired point count Ndesiredfixed . This update follows a volume-based scaling strategy rather than the linear scaling used in LOCUS 2.0 [ 34 ] . (d) Our method adaptively computes the desired point count Ndesired,t based on the scale indicator mˉt and determines the corresponding voxel size dt via proportional-derivative (PD) control with sensitivity-informed gain scheduling. This process corresponds to Fig. 3 (b).
Figure 3 : System overview of GenZ-LIO . (a) In the preprocessing stage, forward propagation uses IMU measurements to propagate the state and covariance, and backward propagation removes motion distortion from the LiDAR scan, yielding the deskewed scan St . (b) The scan St is voxelized using the voxel size from the previous timestep to produce the temporary voxelized scan Vtemp,t . The median range mt of points in Vtemp,t is inserted into a sliding window to compute the spatial-scale indicator mˉt . Based on mˉt , a target number of voxelized points is set as a scale-informed control setpoint, and the voxel size dt is adaptively adjusted via a PD controller with gain scheduling. The updated dt is then used for bi-resolution voxelization of St , yielding Vmerge,t with dt/2 for map integration with reduced discretization error and Vt with dt for state update. (c) The voxelized scan Vt is aligned with the voxel map for a voxel-pruned correspondence search, which avoids unnecessary traversal of neighboring voxels. This process produces the point-to-plane correspondence set Cpl and point-to-point correspondence set Cpo , which are used in the hybrid-metric state update. Finally, the transformed Vmerge,t is integrated into the voxel map. For clarity, the control flow of the scale-aware adaptive voxelization is further illustrated in Fig. 2 (d), where the bi-resolution voxelization step is omitted for comparison with other adaptive voxelization strategies.
Figure 4 : Our scale indicator mˉt mapped onto the estimated trajectory for the Corridor 02 sequence of the SuperLoc [ 56 ] dataset, demonstrating its variation with the scene’s spatial scale.
Figure 5 : Candidate voxel selection based on the region occupied by a query point within its corresponding root voxel. The root voxel is divided into 27 regions, and the occupied region falls into one of four cases: (a) center case, selecting no neighboring voxel; (b) surface case, selecting one surface-sharing neighboring voxel; (c) edge case, selecting three edge-sharing neighboring voxels; and (d) corner case, selecting seven corner-sharing neighboring voxels. The selected neighboring voxels, together with the root voxel, are considered candidate voxels for correspondence search. In Algorithm 2 , the function GetCandidateVoxels selects the candidate voxels based on these sharing relations.
Notation
Explanation
⊞ / ⊟
The encapsulated “boxplus” and “boxminus” operations on the state manifold
W(⋅)
A vector (⋅) in the world frame
L(⋅)
A vector (⋅) in the LiDAR frame
ITL
The extrinsic of the LiDAR frame w.r.t. the IMU frame
WTI
The pose of the IMU frame w.r.t. the world frame
x,x , xˉ
The true state, propagated estimate, and updated estimate
Table 1 : Some important notations in an error-state iterated Kalman filter.
Figure 6 : Definition of the distance dnbr between a query point and selected neighboring voxels. In Fig. 5 , the distance from the query point to a selected neighboring voxel falls into one of three categories: (a) point-to-surface distance (surface-sharing relation); (b) point-to-edge distance (edge-sharing relation); and (c) point-to-corner distance (corner-sharing relation). In Algorithm 2 , the function ComputeDistanceToVoxel determines the distance type based on these sharing relations between the query point’s occupied region and the selected neighboring voxel.
Figure 7 : Overview of the NarrowWide dataset. (a) Data acquisition environment designed to evaluate robotic systems for disaster-response scenarios. (b) Platforms used for data collection: a tracked robot and two handheld devices. These platforms employ different LiDAR sensors to cover diverse sensing configurations, as detailed in Table 2 . (c) Field experiments across different spatial scales using a tracked robot and handheld sensor platforms. The handheld platforms complement the tracked robot experiments by covering confined spaces that are difficult for the tracked robot to access due to its physical size. Details and traversal routes of the individual sequences are provided in Table 3 and Fig. 8 , respectively.
Figure 8 : Mapping results and estimated trajectories of GenZ-LIO on the six sequences of the NarrowWide dataset.
Platform
Tracked robot
Handheld A
Handheld B
LiDAR
Livox MID-70
Velodyne VLP-16
Livox AVIA
FoV (H × V)
70.4∘ circular
360∘×30∘
70.4∘×77.2∘
Max. range [m]
260
100
450
Frequency [Hz]
10
10
10
IMU
VectorNav VN-100
VectorNav VN-100
Bosch BMI088
Frequency [Hz]
200
200
200
Table 2 : Sensor configurations of the three platforms used to collect the NarrowWide dataset. The platforms correspond to those shown in Fig. 7 (b).
Sequence
Platform
Dist. [m]
Duration [s]
Min. width [m]
# of confined– open transitions
Tracked-01
Tracked robot
265.7
621.0
1.0
8
Tracked-02
Tracked robot
269.2
578.2
1.0
8
Handheld-A-01
Handheld A
263.0
521.9
0.3
6
Handheld-A-02
Handheld A
244.9
402.4
0.3
6
Handheld-B-01
Handheld B
408.7
636.0
0.5
10
Handheld-B-02
Handheld B
415.9
529.1
0.5
8
Table 3 : Characteristics of each sequence in the NarrowWide dataset. The minimum width denotes the width of the narrowest space traversed in each sequence, whereas the widest open areas traversed exceed 100m in all sequences. Note that every sequence returns to its starting position, forming a loop.
Scale
Dataset / Sequence
FAST-LIO2 [ 48 ]
Faster-LIO [ 1 ]
AdaLIO [ 26 ]
Point-LIO [ 13 ]
LIO-EKF [ 47 ]
DLIO [ 5 ]
iG-LIO [ 7 ]
PV-LIO [ 32 ] (Baseline)
Baseline w/ adap. vox.
Baseline w/ hybrid-metric
Ours
Confined
SM
Long Corridor
1.77
9.76
1.74
29.10
26.52
2.29
1.55
2.01
1.90
1.90
1.90
Laurel Cavern
3.64
4.22
4.63
5.91
×
0.58
0.36
0.42
0.37
0.39
0.34
SL
Cave 01
×
1.10
×
0.18
×
0.28
0.12
0.15
0.25
0.13
0.12
Cave 02
4.49
0.81
3.98
6.57
×
0.57
0.43
0.43
0.44
0.42
0.44
Cave 04
1.59
×
2.89
0.52
×
6.11
0.17
0.26
0.21
0.22
0.21
H’21
Basement 04
0.63
0.27
0.47
0.50
–
×
0.27
0.13
0.05
0.06
0.05
Table 4 : Absolute translational errors (ATE), reported as RMSE in meters, for each sequence. Due to space limitations, dataset names are abbreviated as follows: SubT-MRS [ 55 ] ( SM ), SuperLoc [ 56 ] ( SL ), 2021 HILTI [ 15 ] ( H’21 ), 2022 HILTI [ 54 ] ( H’22 ), GEODE [ 6 ] ( GD ), M3DGR [ 51 ] ( M3D ), NTU-VIRAL [ 27 ] ( NV ), ENWIDE [ 29 ] ( EW ), and Oxford Spires [ 42 ] ( OS ). The best result is shown in bold . Note that “ × ” denotes a divergent run, defined as an ATE RMSE exceeding 200m , whereas “–” indicates that the method cannot process the LiDAR sensor used in the sequence. The divergence rate is computed as the percentage of divergent runs among the sequences that the method can process. For each sequence, the method with the lowest ATE is ranked first, and the resulting ranks are averaged across sequences. Divergent cases receive the worst rank, and cases marked with “–” are excluded.
Figure 9 : Qualitative comparison on the Exp 16 sequence of the 2022 HILTI [ 54 ] dataset. (a) Onboard views captured in a confined staircase. (b) Divergence of the baseline [ 32 ] in the confined staircase shown in (a). (c) Stable odometry estimation by the baseline with adaptive voxelization in the same staircase. Note that the coordinate frame corresponds to the body frame of the mobile platform. The accumulated map is shown in gray, while the current voxelized scan is shown in blue.
Figure 10 : Qualitative comparison on the Waterways-Short sequence of the GEODE [ 6 ] dataset. (a) Onboard view captured in an open waterway. (b) Pose drift of the baseline [ 32 ] in the open area shown in (a). (c) Stable odometry estimation by the baseline with the hybrid-metric state update in the same open area. Note that the coordinate frame corresponds to the body frame of the mobile platform. The accumulated map is shown in gray, while the point-to-plane constraints are shown in blue and the point-to-point constraints in red.
Figure 11 : Qualitative comparison on the Handheld-A-01 sequence of the NarrowWide dataset. (a) Data acquisition using a handheld device in an extremely confined corner less than 0.5m wide. (b) Divergence of the baseline [ 32 ] in the confined corner shown in (a). (c) Stable odometry estimation by GenZ-LIO in the same corner. This scene corresponds to region C in Fig. 12 . Note that the coordinate frame corresponds to the body frame of the handheld device. The accumulated map is shown in gray, while the point-to-plane constraints are shown in blue and the point-to-point constraints in red.
Figure 12 : Comparison of adaptive voxelization strategies [ 34 , 26 , 9 ] applied to a baseline system [ 32 ] on the Handheld-A-01 sequence of the NarrowWide dataset. (a) Temporal evolution of the proposed scale indicator mˉt . The markers A–F correspond to the regions A–F in Fig. 1 , respectively. (b) Voxel sizes adjusted by each adaptive strategy. (c) Number of points in the voxelized scan obtained using the adjusted voxel size. (d) Relative translational error of the estimated poses. (e) CPU usage, visualized using a sliding-window temporal average with a shaded band indicating one standard deviation to reduce short-term fluctuations. (f) Computation time per frame, visualized using the same smoothing scheme for clarity. The baseline system diverged during the sequence and is therefore plotted only up to the divergence point, which is indicated by a × marker. These results are further summarized in Table 5 .
Method
CPU [%]
Comp. time [ms]
RTE [cm]
ATE [m]
Mean
p95
Mean
p95
RMSE
RMSE
Baseline [ 32 ]
270.12
485.02
16.83
36.70
×
×
+ LOCUS 2.0 [ 34 ]
273.92
477.59
20.98
42.65
1.49
0.29
+ AdaLIO [ 26 ]
294.57
467.87
24.14
50.35
1.56
0.34
+ LIVOX-CAM [ 9 ]
245.93
372.98
19.60
34.59
1.60
0.30
+ Ours
212.23
314.16
16.64
26.24
1.42
0.16
Table 5 : Comparison of computational efficiency and odometry accuracy for different adaptive voxelization strategies evaluated on the Handheld-A-01 sequence of the NarrowWide dataset. CPU usage and computation time are reported using the mean and the 95th percentile (p95). RTE and ATE denote the relative and absolute translational errors, respectively, both reported as RMSE. All adaptive voxelization methods [ 34 , 26 , 9 ] are evaluated within the same baseline framework [ 32 ] . The baseline values shown in gray report CPU usage and computation time only up to the divergence point and are therefore excluded from direct comparison. The symbol “ × ” indicates divergence, and the best performance is highlighted in bold . The temporal behavior of each method over the sequence is further illustrated in Fig. 12 .
Figure 13 : Setpoint tracking performance comparison on the Exp 16 sequence of the 2022 HILTI [ 54 ] dataset. The plots show the temporal evolution of the scale-informed setpoint Ndesired,t and the corresponding voxelized point count Nt for the following ablations: (a) PD controller with fixed gains, (b) ours without scale indicator mˉt , (c) ours without the error terms ∣et∣ and ∣Δet∣ in gain scheduling, and (d) ours with sensitivity-informed gain scheduling. Region A in (b) and region B in (c) correspond to the zoomed regions in (d), highlighting differences in local tracking behavior.
Method
IAE [ ×103 ] ↓
Overshoot ↓
ATE [m] ↓
Linear scaling [ 34 ]
14.03
0.24
0.21
Volume-based scaling [ 9 ]
270.71
0.87
0.34
Fixed gains
53.74
1.63
0.24
Ours w/o scale indicator
13.32
0.16
0.19
Ours w/o error terms
7.95
0.15
0.21
Ours
6.04
0.09
0.13
Table 6 : Ablation study results on voxel size control strategies for tracking a scale-informed setpoint Ndesired,t , evaluated on the Exp 16 sequence of the 2022 HILTI [ 54 ] dataset. IAE denotes the integral of absolute error; see Sec. 6-D . All control methods are evaluated within the same baseline odometry framework. The best performance is highlighted in bold .
Figure 14 : Box plots of the condition number on (a) the Waterways-Short sequence of the GEODE [ 6 ] dataset and (b) the Handheld-B-02 sequence of the NarrowWide dataset, evaluated by applying the proposed modules to the baseline system [ 32 ] . A lower condition number indicates improved numerical stability of the system [ 8 ] . The **** annotations indicate paired t -test results with p<10−4 .
Figure 15 : Mapping result of GenZ-LIO on the Waterways-Short sequence of the GEODE [ 6 ] dataset. Translational and rotational directions that are weakly constrained due to insufficient geometric constraints, and are therefore susceptible to LiDAR degeneracy, are indicated by orange arrows. The visualized coordinate frame corresponds to the robot body frame, and the camera image is included solely for improved scene understanding.
Figure 16 : Average computation time and ATE (RMSE) on the Offroad-04 sequence of the GEODE [ 6 ] dataset under different correspondence search strategies for the following systems: (a) GenZ-LIO and (b) LIO-EKF [ 47 ] .
Stage
Method
Adaptation criterion
Control strategy
Adaptive voxelization
LOCUS 2.0 [ 34 ]
Fixed target number of voxelized points Ndesiredfixed
Linear scaling
AdaLIO [ 26 ]
Fixed threshold for the number of voxelized points τNfixed
Threshold-based switching between two predefined voxel sizes
LIVOX-CAM [ 9 ]
Fixed target number of voxelized points Ndesiredfixed
Volume-based scaling
GenZ-LIO
Scale-informed target number of voxelized points Ndesired,t
PD control with gain scheduling
Stage
Method
Residual type
Residual weighting
State update
PV-LIO [ 32 ]
Point-to-plane [ 36 ] residual
LiDAR measurement noise model [ 50 ]
Table 7 : Comparison between GenZ-LIO and related methods.
Figure 17 : Fixed point-count setpoint analysis under different spatial scales. (a) Distribution of the scale indicator mˉt for each evaluated sequence, where the white diamond and black line indicate the mean and median, respectively. (b) ATE and average computation time obtained by fixing the target point count to Ndesired,t∈{1,000,2,000,3,000,4,000} . The filled markers identify the setpoints with the lowest ATE. (c) Sequence-wise ACE score SNACE , which jointly summarizes ATE and computation time; higher values indicate a better accuracy–efficiency trade-off within the same sequence.
Figure 18 : Comparison of the frame-wise medians of the point-to-plane covariance Rplℓ and the normalized point-to-point covariance Rponorm,ℓ over time on the Tracked-01 sequence of our NarrowWide dataset. The two covariance terms show a persistent numerical scale gap throughout the sequence.
λpo
Stairs
Waterways-Short
Tracked-01
Average rank
0.01
0.25
1.73
0.31
4.00
0.05
0.24
0.69
0.16
1.00
0.1
0.36
0.91
0.17
2.67
0.2
0.31
0.95
0.18
3.00
0.5
0.39
2.42
0.22
5.00
1
0.38
4.10
0.28
5.33
Table 8 : Sensitivity analysis of the point-to-point covariance scaling factor λpo on the Stairs and Waterways-Short sequences of GEODE [ 6 ] and the Tracked-01 sequence of our NarrowWide dataset. ATE is reported as RMSE in meters. The lowest ATE and the lowest average rank are shown in bold .
Figure 19 : Temporal evolution of the frame-wise minimum combined point-to-point covariance Rpocomb,ℓ with and without the discretization-error variance Rdisc . The minimum in each frame is computed over all point-to-point correspondences generated in that frame. Stairs and Waterways-Short are sequences of the GEODE [ 6 ] dataset, and Tracked-01 is a sequence of our NarrowWide dataset.
Method
Stairs
Waterways-Short
Tracked-01
Full Rpocomb,ℓ (Proposed)
0.24
0.69
0.16
Rpocomb,ℓ w/o Rdisc
0.76
1.09
3.35
Table 9 : Ablation study on the combined point-to-point covariance with and without the discretization-error variance Rdisc . Stairs and Waterways-Short are sequences of the GEODE [ 6 ] dataset, and Tracked-01 is a sequence of our NarrowWide dataset. ATE RMSE is reported in meters and the best result on each sequence is shown in bold .
Parameter / Value
Exp 16
Offroad-02
Handheld-A-01
τm
20.0
0.307
0.305
0.221
30.0 ∗
0.196
0.298
0.192
40.0
0.151
0.304
0.205
50.0
0.367
0.301
0.217
p
0.5
0.242
0.300
0.198
1.0
0.224
0.302
0.235
Table 10 : One-at-a-time parameter sensitivity analysis under different spatial-scale conditions. ATE is reported as RMSE in meters. Except for the parameter group indicated in the first column, all parameters are fixed at their default values. “ ∗ ” indicates the default value of each parameter group, and the best result is shown in bold .
Method
NTU-VIRAL [ 27 ]
GrandTour [ 12 ]
Avg. rank
eee_01
sbs_01
tnp_01
KÄB-2
HÖB-2
ALB-2
FAST-LIO2 [ 48 ]
0.89
0.35
0.30
0.07
0.03
0.07
6.00
Faster-LIO [ 1 ]
0.86
0.36
0.24
0.06
0.03
0.06
4.83
AdaLIO [ 26 ]
0.91
0.29
0.20
0.07
0.03
0.07
4.67
Point-LIO [ 13 ]
0.25
0.26
0.24
0.08
0.06
0.09
6.67
LIO-EKF [ 47 ]
0.56
0.28
0.25
0.07
0.68
0.07
6.17
Table 11 : Absolute translational errors on additional structured and unstructured sequences without pronounced spatial-scale variations, reported as RMSE in meters. The best results are shown in bold . Note that “ × ” denotes divergence and receives the worst rank.
LiDAR Inertial Odometry (LIO) is a critical component for many mobile robots that need to navigate without relying on external positioning (e.g., GPS). Platforms that operate autonomously in different environments and with heterogeneous LiDAR sensors require a LIO approach that can adapt to these different scenarios without human intervention. Existing LIO approaches can typically provide reliable and accurate odometry in scenarios with similar environments and sensors when suitably tuned. However, many approaches struggle to retain robust odometry across heterogeneous environments and sensors while using a consistent configuration. This paper presents EllipseLIO, a real-time LIO approach that generalises between scenarios by using methods for LiDAR scan filtering and registration that adapt to the sensor capabilities and environment without requiring scenario-specific tuning. Experiments with EllipseLIO and state-of-the-art LIO approaches on five datasets with diverse and challenging scenarios demonstrate that EllipseLIO is the best-performing approach overall. It achieves a 38% lower odometry error on average than the second-best approach and is the only approach that does not diverge in any experiment. An open-source version of EllipseLIO will be available at github.com/v4rl-ucy/ellipselio.
Rowan Border, Margarita Chli
Vision for Robotics Lab (V4RL), University of Cyprus, Cyprus · ETH Zurich, Switzerland
Reliable odometry is essential for mobile robots as they increasingly enter more challenging environments, which often contain little information to constrain point cloud registration, resulting in degraded LiDAR-Inertial Odometry (LIO) accuracy or even divergence. To address this, we present BIEVR-LIO, a novel approach designed specifically to exploit subtle variations in the available geometry for improved robustness. We propose a high-resolution map representation that stores surfaces as voxel-wise oriented height images. This representation can directly be used for registration without the calculation of intermediate geometric primitives while still supporting efficient updates. Since informative geometry is often sparsely distributed in the environment, we further propose a map-informed point sampling strategy to focus registration on geometrically informative regions, improving robustness in uninformative environments while reducing computational cost compared to global high-resolution sampling. Experiments across multiple sensors, platforms, and environments demonstrate state-of-the-art performance in well-constrained scenes and substantial improvements in challenging scenarios where baseline methods diverge. Additionally, we demonstrate that the fine-grained geometry captured by BIEVR-LIO can be used for downstream tasks such as elevation mapping for robot locomotion.
Patrick Pfreundschuh, Turcan Tuna, Cedric Le Gentil +3
Autonomous Systems Lab, ETH Zürich, Switzerland · Robotic Systems Lab, ETH Zürich, Switzerland · Mobile Robotics Lab, ETH Zürich, Switzerland
Robust and accurate odometry estimation is essential in modern robotics. In environments characterized by highly dynamic motion and sensor noise, odometry estimation becomes increasingly challenging. Autonomous racing combines both factors in an unstructured setting, where minimizing odometry latency is essential for stable closed-loop control. This paper introduces FAR-LIO, a highly optimized CUDA-accelerated LiDAR-inertial odometry framework developed for Fast, Accurate, and Robust performance. Our system leverages a novel CUDA-based voxel hashmap to enable parallelized nearest-neighbor search and efficient map updates. We employ a sparsity-aware Generalized Iterative Closest Point algorithm with adaptive thresholding on top of the CUDA-based voxel hashmap with adaptive density to achieve low-latency without compromising accuracy. An Extended Kalman Filter serves as a robust backend. It utilizes an upsampling and delay compensation strategy to fuse the LiDAR odometry with high-frequency IMU data, thereby ensuring a robust and smooth odometry output. We evaluate FAR-LIO across four different sensor setups, using both public datasets and data from two autonomous racecars driving at speeds of up to 250 km/h. FAR-LIO achieves an average 6.9% reduction in the positional error and 38.4% lower runtime compared to state-of-the-art baselines on target hardware using a single parameter set. This demonstrates its computational efficiency and broad applicability. To build upon our work, our code is available open-source on https://github.com/TUMFTM/FAR-LIO.
Maximilian Leitenstern, Marcel Weinmann, Patrick Haft +3
Institute of Automotive Technology, School of Engineering and Design, Technical University of Munich, Germany · NVIDIA Corp., USA