Globally consistent, real-time state estimation in large-scale, perceptually degraded environments is essential for autonomous vehicles and aerial robots, and requires fusing LiDAR, inertial, and GNSS measurements. Existing fusion methods, however, share a scan-to-map front-end with two failure modes. First, each scan is aligned to an incrementally built map that drifts under degeneracy, and once the estimate diverges the error is irrecoverable. Second, even without divergence, a registration biased by dynamic objects or wrong correspondences is propagated as a single pose constraint with an over-confident covariance, leaving its correspondences unavailable for GNSS to re-weight or relinearize. We propose GLIO2, a tightly-coupled LiDAR-Inertial-GNSS system whose GPU-parallel front-end jointly optimizes scan-to-multiscan LiDAR, IMU pre-integration, and raw GNSS measurements in a single sliding-window factor graph, sustaining real-time operation on edge hardware. A complementary offline back-end reuses the same cached factors to refine the entire trajectory in batch, completing the 30-min, 4.51-km UrbanNav Whampoa sequence in about 24 s. Across three public benchmarks (UrbanNav, MARS-LVIG, M3DGR) and self-collected UAV and vehicle data, GLIO2 attains the best overall accuracy among evaluated systems. On a 5.66-km bridge traversed at up to 96 km/h, where every competing baseline diverges under LiDAR degeneracy, it maintains 1.6 m horizontal accuracy. On an NVIDIA Jetson Orin NX, the full pipeline runs at about 25 Hz (39.60 ms per scan). The source code and datasets will be released.
Figures & tables
Figure 1: GLIO2 in two challenging self-collected scenes, with the point-cloud map (colored by height) projected onto satellite imagery. Top : the 5.66 km Bridge , crossed at up to 96 km/h, where the straight bridge and the heavy moving traffic cause LiDAR degeneracy along the driving direction (front views A, B). Bottom : the 1.44 km dense-canyon Urban route, where high-rise buildings block and reflect the GNSS signals (location C, with fisheye sky view C1 and front view C2). In each scene GLIO2 keeps the global map continuous and aligned with the imagery despite the complementary LiDAR and GNSS degradation.
System
LiDAR Factor
GNSS
Coupling
RT
LIO-SAM [ 3 ]
S2M
GNSS Solution
loose FGO
✓
LIO-Fusion [ 37 ]
S2M
GNSS Solution
loose FGO
✓
He et al. [ 38 ]
S2M
GNSS Solution
loose FGO
×
GIL [ 9 ]
S2M
PPP: PSR+CP
tight EKF
×
P3-LINS [ 31 ]
S2M
PPP: PSR+CP+DOP
tight EKF
×
Li et al. [ 45 ]
S2MS
PPP: PSR+CP
tight MSCKF
✓
Table 1: Summary of the Most Relevant Methods for LiDAR–Inertial–GNSS Integration Systems.
Figure 2: System overview of GLIO2. The front-end IMU-deskews each LiDAR scan and bootstraps by LIO initialization, where the IMU pre-integration and scan-to-multiscan (S2MS) voxelized-GICP (VGICP) LiDAR factors supply motion and local geometry; GNSS is then aligned to the LIO trajectory, and its double-differenced (DD) pseudorange and Doppler factors enter the graph. The sliding-window graph is solved on the CPU at the LiDAR rate with the dense per-point linearization on the GPU: at each linearization point xk , the points are re-associated against the hash-indexed keyframe voxel maps and the block Hessians returned, the CPU solves for the next xk until convergence, and the converged poses update the active keyframe set while out-of-window states are marginalized. The offline back-end reuses the cached relative factors, applies several rounds of GNSS fault detection and de-weighting (FDD), and batch-optimizes the whole trajectory.
Notations
Explanation
E(⋅),w(⋅),l(⋅)
ECEF, local-world, and LiDAR coordinate frames
i(⋅),r(⋅)
IMU and GNSS receiver coordinate frames
ba(⋅)
Variable of frame b represented in frame a
(⋅)k
Physical quantity at discrete time step k
⌊⋅⌋×,(⋅)∧
Skew operator; hat map, R6→se(3)
⊞,⊟,⊕
Manifold plus/minus; orthogonal direct sum
Table 2: Nomenclature of Coordinates and System States
Figure 3: The five coordinate frames of GLIO2: the Earth-centered ECEF frame ( E ), the local-world frame ( w ), and the LiDAR ( l ), IMU ( i ), and GNSS-receiver ( r ) frames on the platform. The LiDAR and IMU share a fixed extrinsic; the local-world frame relates to ECEF through the estimated extrinsic wET .
Figure 4: Double-differenced pseudorange geometry among the rover r , base station b , master satellite m , and secondary satellite i . Double differencing cancels the error terms common to both receivers and both satellites, namely the clock offsets and atmospheric delays.
Figure 5: Block structure of the window information matrix (GNSS in red, LiDAR in blue), contrasting the two LiDAR factors; nodes xi,xj1,xj2,xj3 are the current scan and three keyframes, and each cell is a 6×6 pose subblock. (a) Scan-to-map enters LiDAR as a unary absolute constraint on the diagonal blocks (red-boxed), acting on the same global degrees of freedom as GNSS. (b) Scan-to-multiscan enters LiDAR as relative constraints coupling xi to each keyframe (off-diagonal blocks), so G⊆ker(HL) .
Figure 6: Factor graph of the global batch optimization (Sec. 4-4.9 ): the whole trajectory is re-optimized jointly over all states x0,…,xk under IMU pre-integration factors, the scan-to-multiscan LiDAR between-factors, and the raw double-differenced pseudorange and Doppler factors.
Table 3: Overview of the public and self-collected datasets: platform, sensor suite, reference (ground-truth) source, and the evaluated sequences with their dominant degradation scenario.
Figure 7: Self-collected platforms: (a) , (b) the two UAVs; (c) the ground vehicle (top) and its roof-mounted sensor suite (bottom). Per-platform sensor models and reference sources are listed in Tab. 3 .
Figure 8: Per-sequence GNSS noise and LiDAR registration degeneracy along each timeline, colored green/yellow/red for low/moderate/high severity. The two modalities degrade in complementary sequences: GNSS in Urban and Switch , LiDAR on Bridge and Plain .
Sequence
GLIO2(Online)
GLIO2(Batch)
GLIO(Online) [ 7 ]
GLIO(Batch)
LIGO [ 8 ]
LIO-SAM-GPS [ 3 ]
RTKLIB [ 49 ]
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
UrbanNav-Tst
2.153
3.539
1.064
2.583
5.164
19.559
1.516
3.748
5.935
8.051
8.061
18.930
6.965
11.801
UrbanNav-Whampoa
2.834
7.756
1.955
2.987
7.358
16.627
2.070
5.967
8.506
16.800
10.460
22.710
21.230
41.762
MARS-LVIG-HKairport1
2.545
3.207
0.899
1.701
×
×
×
×
7.002
10.472
42.980
45.355
2.194
9.672
MARS-LVIG-HKisland1
2.260
3.877
2.134
2.473
×
×
×
×
6.999
9.317
13.111
15.530
2.772
13.038
M3DGR-Gra1
0.527
0.534
0.318
0.346
–
–
–
–
0.690
1.334
4.967
9.912
7.974
12.945
Table 4: Horizontal (2D) and 3D ATE RMSE (m) on the public sequences (UrbanNav, MARS-LVIG, M3DGR).
Sequence
GLIO2(Online)
GLIO2(Batch)
GLIO(Online) [ 7 ]
GLIO(Batch)
LIGO [ 8 ]
LIO-SAM-GPS [ 3 ]
RTKLIB [ 49 ]
FAST-LIO2 [ 1 ]
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
Hor.
3D
Urban
1.127
1.978
0.919
1.327
3.529
12.377
1.589
4.492
3.663
12.935
10.468
13.212
15.945
38.877
6.035
14.604
Switch
0.425
0.636
0.213
0.301
2.727
3.303
0.654
0.881
0.734
0.865
1.334
1.442
3.890
11.212
0.564
0.580
Bridge
1.625
3.586
1.438
2.552
×
×
×
×
×
×
×
×
5.634
11.367
×
×
Plain
0.596
0.663
0.142
0.162
×
×
×
×
9.881
9.882
17.580
18.378
0.150
0.289
×
×
Table 5: Horizontal (2D) and 3D ATE RMSE (m) on the self-collected sequences.
Figure 9: Urban sequence: the GLIO2 point-cloud map built online over the full vehicle route ( Start to End ). The four callouts A – D pair each onboard camera view with the local map at a representative location of strong GNSS NLOS and occlusion. The global map stays consistent across the route despite the persistent multipath.
Figure 10: Urban sequence: estimated trajectory (top-down and vertical) against the SPAN-CPT reference, with the inset magnifying the vertical drift over the final GNSS-degraded span ( 250 – 310 s).
Figure 11: Comparison of GLIO2-Tight and GLIO2-Loose on Switch (top) and Urban (bottom). GLIO2-Tight (a, c) keeps a clean, continuous map, whereas GLIO2-Loose (b, d) injects trajectory jumps that tear the map: at the outdoor-to-indoor transition (b) and under tree-shaded GNSS dropouts (d) .
Variant
Switch
Urban
Hor.
3D
Hor.
3D
GLIO2-Tight
0.425
0.636
1.127
1.978
w/o tight GNSS (Loose)
diverges
10.884
12.705
w/o GPU (CPU)
1.570
2.608
2.569
4.064
Table 6: GNSS-coupling and GPU ablation: horizontal and 3D ATE RMSE (m) on Switch and Urban .
Figure 12: Architecture ablation on the Bridge : per-axis position error and the estimator’s reported 1σ for the scan-to-map and S2MS variants in the body frame (lateral x ; forward y , the driving direction; vertical z ), on a log scale; the shaded band marks the LiDAR-degenerate span. Within it the scan-to-map along-track ( y ) error grows by orders of magnitude while its reported 1σ stays millimeter-level, whereas the S2MS fusion keeps the along-track error at the meter level.
Sequence
LiDAR factor
ATE RMSE (m)
2D 1σ (cm)
3D 1σ (cm)
Outcome
Bridge
Scan-to-map
669.6
0.50
0.51
diverges
S2MS (ours)
3.586
1.45
3.31
bounded
Urban
Scan-to-map
11.46
0.22
0.25
drifts
S2MS (ours)
1.978
0.86
1.01
bounded
Table 7: Architecture ablation on the degenerate Bridge and the well-conditioned Urban sequence.
Figure 13: Per-scan runtime distributions across the three LiDAR types on (a) the Jetson Orin NX edge platform and (b) the desktop platform. The dashed line marks the 100 ms real-time budget at the 10 Hz LiDAR rate; for LIGO on the Velodyne in (a) , readouts above 300 ms ( 8.1% of scans) are excluded. In each box the center line is the median and the white diamond ( ◊ ) is the mean, the whiskers span the 1.5× IQR range, and “+” marks outliers.
Per-scan total (ms)
GLIO2 breakdown (ms)
LiDAR
LIGO
FAST-LIO2
GLIO2
Preproc.
Opt. upd.
Residual
Jetson Orin NX
Hesai
152.19
73.56
55.54
15.26
32.06
8.21
Velodyne
185.60
77.72
54.02
11.36
35.24
7.42
Mid-360
123.96
23.28
39.60
5.27
28.04
6.28
Desktop
Table 8: Per-scan runtime (ms) on the Jetson Orin NX and desktop platforms: per-scan total of LIGO, FAST-LIO2, and GLIO2, with the GLIO2 stage breakdown. Fastest total per row in bold.
Figure 14: Back-end batch optimization over the three sequences: 3 D ATE RMSE versus full optimization time (log scale), with bubble size the number of optimized nodes. GLIO2 reaches lower error in far less time while optimizing a larger graph.
Figure 15: Autonomous facade-cleaning UAV. (a) Close view of the UAV holding a 1 – 2.5 m standoff from the glass curtain wall (annotated distances are UAV-to-wall). (b) Online GLIO2 LiDAR point-cloud map of the repetitive facade, whose self-similar planar structure causes LiDAR degeneracy along the wall, with the estimated UAV pose. (c) The 110 m high-rise during the cleaning operation, with the UAV working along the facade.
Figure 16: Guide-dog navigation for the visually impaired. (a) Online GLIO2 point-cloud map. (b) Route on Google Maps imagery. (a1) Point cloud and (a2) real-world view of the LiDAR-degenerate section; (a3) real-world view and (a4) point cloud of the GNSS-denied section.
Robust state estimation and mapping in long-term, large-scale, and highly dynamic environments remains a key challenge in robotics. Existing LiDAR-Inertial-Visual Odometry (LIVO) systems achieve strong local accuracy but suffer from accumulated drift over long distances and may fail in geometrically degraded or textureless scenes. Meanwhile, GNSS-aided fusion frameworks often rely on LiDAR or visual odometry for state prediction and outlier rejection, making them vulnerable when odometry degenerates. To address these limitations, we propose a tightly coupled LiDAR-Inertial-Visual-GNSS fusion framework based on an Error-State Iterated Kalman Filter. An online spatiotemporal alignment module using Dynamic Time Warping is introduced for highly dynamic conditions. To better exploit GNSS precision, we develop observation models based on Doppler shifts and fixed-anchor Time-Differenced Carrier Phase, providing millimeter-level relative constraints without augmenting historical anchor states. We further design a degeneracy-aware dual-mode outlier rejection strategy that switches between LIVO-prior-guided rejection and GNSS-aided recovery according to the LIVO degeneracy level. Experiments on the public M3DGR dataset and a custom 20~m/s fixed-wing UAV dataset demonstrate that our system reduces accumulated drift and map ghosting, outperforming state-of-the-art methods in accuracy and robustness.
Zhiyu Chen, Chunran Zheng, Jiayu Wen +4
College of Mechatronics and Control Engineering, Shenzhen University, Shenzhen 518060, China · Department of Mechanical Engineering, The University of Hong Kong, Hong Kong SAR, China · College of Automation, Harbin Engineering University, Harbin 150001, China
This paper introduces 2Fast-2Lamaa, a lidar-inertial state estimation framework for odometry, mapping, and localization. Its first key component is the optimization-based undistortion of lidar scans, which uses continuous IMU preintegration to model the system's pose at every lidar point timestamp. The continuous trajectory over 100-200ms is parameterized only by the initial scan conditions (linear velocity and gravity orientation) and IMU biases, yielding eleven state variables. These are estimated by minimizing point-to-line and point-to-plane distances between lidar-extracted features without relying on previous estimates, resulting in a prior-less motion-distortion correction strategy. Because the method performs local state estimation, it directly provides scan-to-scan odometry. To maintain geometric consistency over longer periods, undistorted scans are used for scan-to-map registration. The map representation employs Gaussian Processes to form a continuous distance field, enabling point-to-surface distance queries anywhere in space. Poses of the undistorted scans are refined by minimizing these distances through non-linear least-squares optimization. For odometry and mapping, the map is built incrementally in real time; for pure localization, existing maps are reused. The incremental map construction also includes mechanisms for removing dynamic objects. We benchmark 2Fast-2Lamaa on over 750km of public and self-collected datasets from both automotive and handheld systems. The framework achieves state-of-the-art performance across diverse and challenging scenarios, reaching odometry and localization errors as low as 0.22% and 0.06 m, respectively. The real-time implementation is publicly available at https://github.com/clegenti/2fast2lamaa.
Cedric Le Gentil, Raphael Falque, Daniil Lisus +1
Robotics Institute, University of Toronto Institute for Aerospace Studies, Canada · Robotics Institute, University of Technology Sydney, Australia
We present a real-time LiDAR-based framework for Gaussian Splatting SLAM that tightly couples fast G-ICP registration with spherical rasterization-based dense mapping for large-scale sequences. Leveraging LiDAR geometry rather than appearance, we reuse tracking-estimated local covariances to initialize Gaussians with range-aware scales and to derive surface normals for geometry-aware map optimization. We further introduce a covariance-derived geometry score that measures local complexity and drives pruning in planar regions and selective densification in structurally rich areas, while optimized Gaussians and LiDAR-specific confidence cues are fed back to improve tracking robustness. On the Newer College dataset, our method achieves an F-score of 86.78% using purely online trajectories at real-time speed (>20 FPS), and additional experiments on other datasets confirm its stability and scalability.