LBDU-VIO: Learned Bias Dynamics and Uncertainty for Visual-Inertial Odometry with Unreliable Vision
Authors: Qizhi Guo, Junning Lyu, Defu Lin, Shaoming He
Organizations: School of Aerospace Engineering and the Beijing Key Laboratory of UAV Autonomous Control, Beijing Institute of Technology, Beijing 100081, China
Visual-inertial odometry (VIO) for aerial robots relies on high rate inertial measurement unit (IMU) propagation between visual updates. However, conventional multi state constraint Kalman filters (MSCKFs) use random walk bias assumptions and fixed noise parameters, which can limit robustness when visual information is unreliable. To address this problem, we propose LBDU-VIO, a learning-augmented MSCKF with learned continuous time bias dynamics and an IMU uncertainty model. A neural ordinary differential equation (ODE) models continuous time bias dynamics to propagate the filter's bias states, replacing their random walk model. The IMU uncertainty model predicts motion adaptive measurement noise covariances for covariance propagation. Both models are trained with pose supervision without direct labels. Experiments on real world EuRoC and TUM-VI benchmarks show lower errors than representative visual-inertial baselines, including a 25.1% reduction in mean relative position error compared with S-MSCKF on EuRoC sequences with 10s visual outage.
Figures & tables
Fig. 1: Conceptual illustration of VIO under unreliable vision. The proposed LBDU-VIO aims to reduce drift during visual degradation and visual recovery after temporary tracking failure.
Fig. 2: LBDU-VIO system pipeline. A learning based multi state constraint Kalman filter that combines learned IMU bias dynamics and uncertainty with visual updates. Instead of modeling IMU biases as random walks and using fixed noise parameters, LBDU-VIO corrects raw IMU measurements with bias estimates from a neural ODE and uses measurement noise covariances predicted by an uncertainty model in covariance propagation.
Fig. 3: Schematic comparison of bias representations. Learned continuous time dynamics describe bias evolution within each window, whereas a single estimate per window gives a piecewise constant bias.
Dataset
IMU
Rate
Env.
Train / Test
EuRoC MAV
ADIS16448
200 Hz
Indoor
6 / 5
TUM-VI
BMI160
200 Hz
In/Outdoor
3 / 3
TABLE I: Datasets and train–test splits
Parameter
Symbol
EuRoC
TUM-VI
Unit
Gyro. white noise
σg
1.70×10−4
1.60×10−4
rad/s/ Hz
Accel. white noise
σa
2.00×10−3
2.80×10−3
m/s 2 / Hz
Gyro. bias diffusion
σbg
1.94×10−5
2.20×10−5
rad/s 2 / Hz
Accel. bias diffusion
σba
3.00×10−3
8.60×10−4
m/s 3 / Hz
TABLE II: IMU noise density parameters for EuRoC and TUM-VI
S-MSCKF [ 40 ]
MSCEqF [ 41 ]
VIO-IPNet [ 28 ]
Brossard [ 11 ]
LBDU-VIO
Dataset
Sequence
APE
AOE
APE
AOE
APE
AOE
APE
AOE
APE
AOE
EuRoC
MH_02
0.141
1.860
0.639
3.117
0.166
1.721
0.196
2.334
0.232
2.221
MH_04
0.355
2.243
0.499
1.818
0.325
1.790
0.300
1.701
0.215
1.194
V1_01
0.090
1.276
0.132
6.185
0.086
0.938
0.098
1.151
0.055
0.470
V1_03
0.247
6.911
0.167
2.106
0.095
2.514
0.093
2.299
0.089
3.155
V2_02
0.163
3.097
0.208
3.149
0.098
1.839
0.129
1.955
0.076
1.412
TABLE III: VIO accuracy on the EuRoC and TUM-VI test sequences under nominal visual conditions
Fig. 4: Relative translation drift on the five EuRoC test sequences.
Outage
RPE (m)
ROE ( ∘ )
Best Baseline
Proposed
Reduction
Best Baseline
Proposed
Reduction
1 s
0.052
0.044
15.4%
0.304
0.275
9.5%
2 s
0.053
0.046
13.2%
0.308
0.273
11.4%
3 s
0.060
0.044
26.7%
0.308
0.254
17.5%
5 s
0.079
0.058
26.6%
0.306
0.259
15.4%
10 s
0.187
0.140
25.1%
0.320
0.306
4.4%
TABLE IV: Mean RPE and ROE across the five EuRoC test sequences at each visual-outage duration. The best baseline is selected separately for each metric and duration. Best values are shown in bold.
Fig. 5: Visual outage and recovery on EuRoC MH_04_difficult sequence. (a) Trajectory APE after global SE(3) alignment. (b) Recovery RPE denotes 1 s relative translation RMSE over the first 10 s after images resume. (c) Position error with pre outage alignment. Shading marks the 50 - 60 s outage and red dotted lines indicate local recovery near 62 s.
Fig. 6: Trajectories on TUM-VI room4 with a 10 s visual outage. (a) XY projections over 20 - 110 s, with planar RMSE reported above each plot. Thick segments show the outage, and connected squares mark XY errors at 60 s. (b) Three dimensional trajectories during the outage. Circles and squares mark 50 and 60 s. Each method shows the run with median APE among three repeats.
Method
AOE RMSE ( ∘ )
APE RMSE (m)
MSCEqF
2.619
0.175
Brossard
1.779
0.194
VIO-IPNet
2.181
0.152
LBDU-VIO
1.378
0.101
TABLE V: Estimation accuracy under intermittent visual distortion on EuRoC V2_02_medium. Values are median RMSEs and bold indicates the lowest value in each column.
Fig. 7: Orientation (a) and position (b) errors under intermittent visual distortion on EuRoC V2_02_medium. The upper track shows the imposed horizontal image shift. Lines show the median error at each timestamp, and shaded bands span the minimum and maximum errors across three runs.
Fig. 8: Orientation errors during open loop gyroscope integration on the TUM-VI test sequences: (a) room2, (b) room4, and (c) room6. Larger Raw IMU errors are clipped at the upper boundary, and the upward arrow shows that this trace continues beyond the displayed range.
Method
MH02
MH04
V103
V202
V101
Mean
APE / RPE (m)
Raw IMU + Fixed Noise
0.153 / 0.087
0.306 / 0.104
0.085 / 0.075
0.606 / 0.158
0.097 / 0.093
0.249 / 0.103
Learned Bias + Fixed Noise
0.145 / 0.055
0.284 / 0.078
0.085 / 0.059
0.150 / 0.057
0.127 / 0.094
0.158 / 0.069
Raw IMU + uncertainty model
0.160 / 0.088
0.296 / 0.107
0.095 / 0.074
0.557 / 0.133
0.090 / 0.089
0.240 / 0.098
Learned Bias + uncertainty model
0.107 / 0.050
0.227 / 0.060
0.077 / 0.050
0.088 / 0.039
0.054 / 0.024
0.111 / 0.045
TABLE VI: Component ablation under a three-second visual outage on the EuRoC test sequences. Values are reported as APE/RPE in meters. Lower values indicate better accuracy, and the best results are shown in bold.
Visual inertial odometry (VIO) is essential for accurate 6-DoF motion estimation in mobile robotic systems. Recent learning-based VIO methods have shown promising progress, but they often rely on unified visual--inertial representations and a single temporal model for full-pose estimation, limiting their ability to capture the heterogeneous dynamics of rotation and translation. Moreover, monocular visual features often lack explicit geometric structure, while raw inertial encoding leaves the underlying rotational kinematics implicit, weakening the rotation-related cues in IMU features. To address these issues, we propose DB-VIO, a dual-branch visual inertial odometry framework with enhanced visual--inertial representation. DB-VIO incorporates depth cues to improve monocular visual perception, injects an explicit integrated-attitude prior to strengthen rotation-aware inertial representation, and decouples pose estimation into dedicated rotational and translational branches for motion-specific temporal modeling. Experiments on autonomous driving and aerial robot benchmarks show that DB-VIO achieves state-of-the-art performance, improving the corresponding baselines by 20% on KITTI and 33% on EuRoC. Notably, under the more agile motion patterns of EuRoC, DB-VIO improves the rotational metric by 65.7% over prior methods. These results demonstrate the effectiveness and generalization of DB-VIO across different platforms and motion scenarios.
Ziyu Wan, Lin Zhao
Department of Electrical and Computer Engineering, National University of Singapore, Singapore.
This work introduces a hybrid deep learning approach integrated with an Unscented Kalman Filter (UKF) to enhance pose estimation accuracy in Visual-Inertial Odometry (VIO) for autonomous navigation. The proposed model employs a Vision Transformer (ViT) network to effectively capture temporal dependencies from inertial measurement unit (IMU) data and utilizes a Multiscale Convolutional Neural Network (MCNN) to learn optical flow-based motion cues from visual data. An adaptive sensor fusion module dynamically weights IMU and visual features by leveraging estimated uncertainty, thus improving robustness in diverse and challenging environmental conditions. Additionally, a novel uncertainty-aware loss function is proposed to explicitly incorporate prediction uncertainty into the learning process, enabling robust and accurate navigation under noisy, incomplete, or unreliable sensor inputs. Comprehensive evaluations of the KITTI dataset demonstrate that the proposed method significantly outperforms baseline approaches, achieving superior performance in terms of Absolute Trajectory Error (ATE) and Relative Pose Error (RPE). The lightweight and computationally efficient model processes data at 155 FPS on an NVIDIA A100 GPU, making it highly suitable for deployment in resource-constrained autonomous systems.
Simegnew Yihunie Alaba, Yuichi Motai
Department of Electrical and Computer Engineering, College of Engineering, Virginia Commonwealth University, Richmond, VA 23220, USA
Learning is increasingly introduced into visual-inertial odometry (VIO), ranging from learned feature front-ends to learning-dominant motion and geometry estimation. However, learning more of the pipeline does not necessarily improve robustness when deployment conditions differ from the training distribution. This work asks whether robust VIO under distribution shift truly requires deeper learned estimation, or whether learning can be confined to visual measurement generation. We propose a minimal-learning stereo VIO framework in which SEA-RAFT is used only to propose dense stereo correspondences and predict their uncertainty, while temporal tracking, geometric verification, and state estimation remain explicit. Dense flow is sampled at sparse feature locations, filtered using predicted uncertainty and stereo epipolar consistency, and incorporated into a sliding-window stereo-inertial estimator through uncertainty-weighted reprojection factors. The same uncertainty is further propagated through stereo triangulation for downstream anisotropic 3D Gaussian mapping. Experiments on EuRoC, VIODE, and 4Seasons demonstrate accurate and stable estimation under motion blur, dynamic scenes, illumination changes, and large indoor-to-outdoor distribution shifts. Ablations show that learned flow alone is insufficient: the gains arise from combining learned correspondence proposals with geometric verification and uncertainty-aware weighting. These results suggest that, for OOD-robust VIO, carefully integrated learned visual measurements can be more effective than learning a larger fraction of the estimation pipeline. Code and configs for the benchmark will be open-source upon acceptance. A supplementary video is available at https://drive.google.com/file/d/1EVRhOkhanmNXHbQS1Vr80FoEIAYOYOV2/view
Yangyang Ning, Shu Liang, Quanbo Ge +3
Department of Control Science and Engineering, Tongji University, Shanghai 201804, China