Observability Analysis and Online Calibration of Visual-Inertial-Wheel Odometry for 4WIS4WID Mobile Robots
Authors: Branimir Ćaran, Vladimir Milić, Bojan Šekoranja, Bojan Jerbić
Organizations: Faculty of Mechanical Engineering and Naval Architecture, University of Zagreb, Zagreb, 10000, Croatia · Croatian Academy of Sciences and Arts, Zagreb, 10000, Croatia
In this paper, we present a visual-inertial-wheel odometry (VIWO) framework with online calibration for four-wheel independently steered and driven (4WIS4WID) mobile robots. We derive a 2D odometry model directly from the four driving velocities and steering angles, using both the longitudinal rolling constraints and the lateral no-slip constraints of all wheels. A preintegration model and analytical Jacobians are developed for efficient filtering and calibration. An observability analysis of the linearized VIWO system shows that a drive-only model makes all steering offsets unobservable, whereas the proposed redundant model restores their observability. The analysis also identifies four standard VINS unobservable directions and three additional directions associated with the arbitrary placement of the odometry reference frame. Furthermore, we characterize several degenerate motions, including zero yaw rate, constant steering, and a non-rolling wheel, and derive the corresponding excitation conditions for the thirteen wheel intrinsics that remain after fixing the odometry frame reference. The performance of the proposed system has been demonstrated in both simulation and real-world experiments on a 4WIS4WID mobile robot.
Figures & tables
Fig. 1: 4WIS4WID mobile robot where each wheel i is independently steered by δi and driven at ωdi , so the ICR may be placed anywhere in the plane.
Motion
Unobservable
General motion
4 inertial + 3 gauge
Zero yaw rate ( ω≡0 )
xw , yw
Constant steering configuration
2 per wheel (8 total)
Wheel j not rolling
rj , δoj
Pure translation
OpI
One axis rotation
OpI along axis
TABLE I: Degenerate motions and the corresponding unobservable calibration parameters.
Fig. 2: Estimation error (solid) and 3σ bound (dotted) for six Monte-Carlo runs along the ICR-sweep trajectory. Shown are the IMU-odometry extrinsics and representative wheel intrinsics ( r1 , δo2 , xw2 ).
Configuration
RPE 10 [deg]
RPE 10 [m]
RPE 25 [deg]
RPE 25 [m]
NEES
True init., calib. ON
0.014
0.0045
0.021
0.0055
3.56
True init., calib. OFF
0.050
0.0071
0.081
0.0111
10.39
Bad init., calib. ON
0.014
0.0046
0.018
0.0060
3.72
Bad init., calib. OFF
0.602
20.9823
1.465
47.4002
170.75
TABLE II: Relative pose error (RPE) and mean NEES over 20 Monte-Carlo runs.
Fig. 3: Experimental results
Parameter
Before
After
wheel radii r [mm]
[25.6, 25.6, 25.6, 25.6]
[26.3, 26.3, 28.1, 26.8]
wheel positions xw [m]
[-0.1125, -0.1125, 0.1125]
[-0.1029, -0.1029, 0.1454]
wheel positions yw [m]
[0.1125, -0.1125, -0.1125]
[0.1267, -0.1216, -0.1216]
steering offsets δo [rad]
[0.0, 0.0, 0.0]
[0.0168, -0.0098, -0.0633]
Ext. Pos [m]
[0.1457, 0.0279, 0.0576]
[0.1696, 0.0192, 0.0244]
Ext. Ori [rad]
[1.2085, 1.2093, 1.2091]
[1.2498, 1.2060, 1.2206]
TABLE III: Values of the calibration parameters before and after online calibration for real world experiment
Visual-Inertial Odometry(VIO), which is critical to mobile robot navigation, uses cameras with a large number of pixels. Capturing and processing camera images requires significant resources. This work presents a minimalist approach to planar odometry, demonstrating that just four visual measurements and an IMU can provide robust motion estimation for differential-drive robots. Our key insight is that four downward-facing photodiodes that sense the world through optical Gabor masks produce signals that encode speed. Based on this, we jointly optimize the mask parameters alongside a Temporal Convolutional Network (TCN) using a physically-grounded simulator. The resulting model decodes speed from just the four measurements produced by the photodiodes. Pairing these estimates with the angular speed from an IMU yields a continuous planar trajectory. We validate our approach with a prototype sensor mounted on a differential drive robot. Across diverse indoor and outdoor terrains, our system closely tracks the reference ground truth without any real-world fine-tuning. Our work shows that minimalist sensing enables efficient and accurate planar odometry.
Francesco Pasti, Jeremy Klotz, Nicola Bellotto +1
Department of Information Engineering, University of Padua, Padua, Italy · Computer Science Department, Columbia University, New York, NY, USA
Monocular visual-inertial odometry (VIO) cannot recover metric scale from vision alone; scale must be resolved through inertial measurements. We present a trajectory-dependent observability analysis showing that translational acceleration, produced by curvature, not constant-speed straight-line travel, is the fundamental source that couples scale to the inertial state. This relationship is formalized through the gravity-acceleration asymmetry in the IMU model, from which we derive rank conditions on the observability matrix and propose a lightweight excitation metric computable from raw IMU data. Controlled experiments on a differential-drive robot with a monocular camera and consumer-grade IMU validate the theory, with straight-line motion yielding 9.2% scale error, circular motion 6.4%, and figure-eight motion 4.8%, with excitation spanning four orders of magnitude. These results establish trajectory design as a practical mechanism for improving metric scale recovery.
Hadush Hailu, Bruk Gebregziabher, Siddhartha Gudipudi +1
Computer Science, Maharishi International University, Iowa, USA · Electrical Engineering and Computing, University of Zagreb, Zagreb, Croatia · Computer Science, Iowa State University, Iowa, USA +1
Deformable scenes violate the rigidity assumptions underpinning classical visual--inertial odometry (VIO), often leading to over-fitting to local non-rigid motion or to severe camera pose drift when deformation dominates visual parallax. In this paper, we introduce DefVINS, the first visual-inertial odometry pipeline designed to operate in deformable environments. Our approach models the odometry state by decomposing it into a rigid, IMU-anchored component and a non-rigid scene warp represented by an embedded deformation graph. As a second contribution, we present VIMandala, the first benchmark containing real images and ground-truth camera poses for visual-inertial odometry in deformable scenes. In addition, we augment the synthetic Drunkard's benchmark with simulated inertial measurements to further evaluate our pipeline under controlled conditions. We also provide an observability analysis of the visual-inertial deformable odometry problem, characterizing how inertial measurements constrain camera motion and render otherwise unobservable modes identifiable in the presence of deformation. This analysis motivates the use of IMU anchoring and leads to a conditioning-based activation strategy that avoids ill-posed updates under poor excitation. Experimental results on both the synthetic Drunkard's and our real VIMandala benchmarks show that DefVINS outperforms rigid visual--inertial and non-rigid visual odometry baselines. Our source code and data will be released upon acceptance.