Sparse inertial pose estimation promises camera-free motion capture from consumer devices, but consumer sensors are unreliable: firmware-fused orientations are biased, mounting varies between sessions, and streams drift or drop out. On a new 35-take single-subject benchmark pairing an earbud head inertial measurement unit (IMU) with two smart-insole foot IMUs (SAM-3D-Body pseudo-ground-truth labels), we show the reliability problem is channel-level: a channel ablation isolates foot acceleration as the most informative input (66.6 mm vs. 79.0 mm head-only) and the firmware-fused foot orientation as the liability that destroys the gain. We therefore let the model learn how much to trust each channel of each stream: one temporal gate per stream per channel block, trained with an auxiliary reliability objective on synthetically corrupted pretraining data. The channel-gated model is the most accurate of our learned fusion arms on clean data (69.4 mm vs. 83.7 static, 86.6 ungated) and under every simulated fault (bias in training; drift, dropout eval-only); its gates suppress the natively biased foot-orientation channels on clean real data without test-time supervision and flag dropout bursts at 0.92-0.999 AUROC. Two contrasts: dropping a channel known a priori to fail is flat across foot faults but collapses when an unanticipated stream fails (head dropout: 92.9 vs. 79.3 mm); and a fine-tuned HMD-Poser is more accurate on clean data (64.4 mm) and nominally under drift, with no significant paired difference under bias or dropout, but a larger worst-case degradation from clean (+16.1 vs. +3.5 mm, single seed). Learning to gate reliability instead of sensor count is the lever for deployable sparse inertial capture. Code is available at https://github.com/ZhilinGuo/reliability-gated-imu-fusion.
Figures & tables
Figure 1 : Overview. (a) Training-time capture: four RGB-D views (labels only), an earbud head IMU, and two smart-insole foot IMUs, with no pelvis sensor and no foot gyroscope. (b) Channel-reliability gating: one temporal gate per stream per channel block learns how much, and when, to trust orientation vs. acceleration, supervised only by synthetic corruption in pretraining. (c) Held-out prediction (red) vs. pseudo-GT (grey): foot acceleration alone reaches 66.6 mm (79.0 mm head-only), the channel gate is the most accurate of our learned fusion arms in every fault regime, and the gates flag dropout bursts at 0.92–0.999 AUROC without test-time supervision.
Figure 2 : Reliability-gated fusion. Per-stream stems (head: earbud; feet: smart insoles; split per channel block) feed scalar temporal gates gi,b(t) , a shared Conv1d ( k=9 ) → ReLU → sigmoid producing one scalar per stream × block. Gated features are concatenated and projected into a two-layer bidirectional LSTM trunk ( 2×512 /direction) regressing 6D rotations for nine lower-body joints, then forward kinematics with fixed SMPL bone lengths and a contact head. Stage 1: pretrain on synthetic AMASS with foot corruption ( p=0.5 ) and a gate BCE objective. Stage 2: fine-tune on real pseudo-GT with an aligned-FK ℓ1 loss (gate loss off).
Figure 3 : Channel-gated predictions (red) vs. SAM-3D-Body pseudo-GT (grey) on a held-out vertical-motion take (squats and lunges), three frames. The full three-class montage and a head-only contrast are in supp. S4.
Method
sim-MPJPE ↓
Rmotion2↑
distal sim ↓
distal R2↑
MPJVE ↓
PA-MPJPE ↓
Constant mean-pose prior
105.5
0.03
141.2
-0.07
220
54.8
Zero-input control (run)
116.4
-0.40
141.6
-0.09
221
63.9
IMUPoser-adapted
Synthetic only (zero-shot)
316.5
-4.93
441.0
-6.84
491
218.3
No calibration, head+feet (run)
103.2
-0.15
124.0
0.16
317
60.6
Head IMU only (run)
79.0
0.06
87.0
0.49
253
57.3
Table 1 : Full 35-take benchmark. Position errors (sim-MPJPE, distal, PA-MPJPE) in mm; distal averages knees, ankles, and feet; MPJVE in mm/s; “run”/“seq”: leave-one-run-out / leave-one-motion-out. Full foot streams cost nothing detectable; removing mounting calibration degrades accuracy badly.
arm
clean
bias σ=10∘
bias σ=20∘
drift
dropout
no gate (frozen g=1 )
86.6 ± 1.8
93.5 ± 0.4
100.5 ± 1.4
87.6 ± 1.1
94.1 ± 1.1
stream-dropout training
88.7 ± 0.3
92.7 ± 0.7
98.1 ± 2.6
89.2 ± 0.6
92.8 ± 1.1
static learned weights
83.7 ± 0.9
87.2 ± 0.2
93.3 ± 0.9 †
83.8 ± 1.1
85.4 ± 0.8 ‡
cross-stream attention
87.3 ± 0.6
—
95.6 ± 0.1
87.6 ± 0.5
89.1 ± 0.4
per-frame gate, no aux.
89.8 ± 3.7
89.5 ± 0.8
96.5 ± 1.2
89.5 ± 3.8
91.4 ± 3.4
per-frame gate
83.9 ± 0.7
88.9 ± 0.4
94.3 ± 0.7 †
84.7 ± 0.8
86.3 ± 0.7 ‡
Table 2 : Gated-fusion matrix (sim-MPJPE mm, run split, three-seed mean ± seed s.d., concat-fusion trunk). Bias is train+test; drift and dropout are eval-only. Bold: best per column (oracle excluded). † / ‡ : better than no gate at Holm-corrected p<0.05 / p<10−6 (paired per-take Wilcoxon, seed-averaged takes, n=35 ; family of 15: 8 at p<0.05 , 5 at p<10−6 ; clean and bias- σ=10∘ columns exploratory; full table: supp. S4). The oracle overrides the per-frame gate arm with whole-take hard masks; attention’s bias- σ=10∘ cell was not run.
Figure 4 : Channel gates flag native bias and burst faults without test-time supervision. (a) Mean gate per stream × channel block across regimes: foot orientation gates are partially closed on clean data (native bias) and close further only when bias enters fine-tuning; foot acceleration gates close only under dropout. (b) On a dropout take, foot acceleration gates close inside corrupted bursts (shaded) and reopen.
model
clean
bias
drift
dropout
worst-case
HMD-Poser (fine-tuned)
64.4
80.4
67.8
74.2
+25%
HMD-Poser (ft, corrupt-aug)
76.2
82.1
78.8
78.1
+8%
MobilePoser (zero-shot)
125.3
127.6
124.7
125.2
+2%
no gate
85.0
100.0
86.4
93.1
+17.6%
static weights
83.4
87.0
83.3
85.1
+4.3%
temporal gate
86.5
90.7
86.9
88.6
+4.9%
Table 3 : External baselines under eval-only foot corruption (sim-MPJPE mm, run split, bias σ=30∘ ; single seed s0). The channel gate is the most accurate of our learned arms under every fault, at 4.6× HMD-Poser’s CPU throughput with 37% fewer parameters (supp. S7). Bold: best of our learned arms per sim-MPJPE column.
Consumer earbuds already stream inertial motion data from the head, one of the most widely worn sensor locations on the body. We ask how much of the 3D body pose a single such head IMU can recover, and whether adding more consumer sensors actually helps. We build a multimodal capture pipeline that records four-view RGB-D video together with an AirPods head IMU and two Striv insole IMUs, synchronize the streams post-hoc, and generate pseudo-ground-truth with SAM 3D Body, yielding a 35-take single-subject benchmark spanning gait, turning, vertical, everyday, and clinically inspired motions. Adapting two recurrent model families (IMUPoser and MobilePoser), we show that one head IMU recovers lower-body pose at 79.0 mm rigid-MPJPE and per-foot ground contact at 0.809 macro-F1, and that a causal variant retains most of this accuracy at streaming latency. In paired per-take significance tests across both families, adding the consumer foot IMUs never significantly improves pose and significantly degrades it in two of four model-split combinations; a mounting-bias probe and feet-only ablation identify insole orientation quality, not foot placement, as the mechanism. Extending the output to a 20-joint full-body skeleton maps the boundary: gross distal-arm motion is partially recoverable from the head alone, proximal upper-body pose is not, and staged fine-tuning recovers the leg accuracy that naive joint training sacrifices to multi-task dilution. For learned pose from consumer wearables, sensor reliability, not sensor count, is the binding constraint here. For the devices tested, the earbud is its sweet spot. Code is available at https://github.com/ZhilinGuo/one-sensor-whole-body.
Zhilin Guo, Boqiao Zhang, Oszkár Urbán +8
University of Cambridge United Kingdom · Chalmers University of Technology Sweden
Methods using inertial measurement units (IMUs) provide a wearable alternative to camera-based motion capture. To mitigate drift from inertial signals, recent sparse inertial pose estimators integrate inter-sensor distances measured by ultra-wideband (UWB) ranging. So far, UWB distances have only been used as an additional input feature, ignoring the physical constraints they impose on sensor positions. However, these distances can also be used to reconstruct the underlying 3D sensor layout, which in turn provides more informative input for pose reconstruction. We propose Ultra Diffusion Poser, a diffusion model that explicitly models these geometric constraints. It includes a Spatial Layout Module that analytically reconstructs the 3D sensor positions from UWB measurements. These sensor positions are used alongside IMU signals and UWB distances as a conditioning signal during diffusion. Still, network predictions can violate inter-sensor distance measurements. To address this, we introduce UWB-Diffusion Guidance, which encourages alignment between predicted poses and measured distances during diffusion sampling. Together, these contributions enable our model to achieve state-of-the-art performance, reducing joint position error by up to 22% over prior work.
Dominik Hollidt, Tommaso Bendinelli, Christian Holz
Department of Computer Science, ETH Zurich, Switzerland
Inertial odometry (IO) using only Inertial Measurement Units (IMUs) provides a lightweight solution for human motion tracking in augmented reality (AR) and wearable devices. Recent learning-based IO methods have improved the generalizability of inertial localization through large-scale pretraining on human motion datasets. However, these approaches remain prone to drift and noise because they do not explicitly capture human motion dynamics, especially on daily activity datasets such as Nymeria. In this work, we propose to ground inertial odometry in human kinematics through a learned IMU-inferred pose prior, which promotes physically consistent motion constraints. We integrate this pose prior into existing IO architectures and reduce positional drift by up to 36% on the challenging Nymeria dataset, which is 5x larger than datasets used in prior work. We further improve long-term performance with a sensor-fusion framework that incorporates auxiliary signals from lightweight sensors already available on commercial AR glasses, including magnetometers, barometers, and secondary IMUs. With this fusion strategy, positional drift is reduced by up to 42%, improving robustness and generalization across diverse motion conditions. Together, our results introduce a new paradigm for inertial and lightweight odometry by unifying human motion kinematics with multimodal sensing, setting a new benchmark for accurate and robust camera-less human tracking. Our website is available at https://spice-lab.org/projects/MARIO/.