Image-aligned metric demonstrations often require dedicated tracking hardware and synchronization across devices. We present MonoEgo, a capture system that replaces active wrist instrumentation with offline monocular reconstruction. One 90-FPS global-shutter camera observes calibrated passive wrist constellations, sparse workstation anchors, and the scene on a shared image clock. MonoTag SLAM combines marker corners with ORB geometry and uses visual evidence to reject ambiguous planar-marker poses. Its metric Atlas supports interval scale re-anchoring, verified map merging, and retrospective localization of earlier frames supported by the final map. Camera and wrist-constellation outputs retain validity and map provenance, and unsupported motion is left missing. Experiments show metric tracking beyond continuous anchor visibility, reconnection of supported map components, and recovery of some missing camera poses. Comparisons against a multisensor camera reference and separate stationary-constellation tests characterize trajectory agreement and precision while revealing incomplete coverage and residual geometric uncertainty. The results indicate that passive fixtures and offline reconstruction can reduce capture-side requirements. Dynamic accuracy, deployment, and downstream policy benefits require further study.
Figures & tables
Fig. 1: MonoEgo capture and offline metric reconstruction. Hardware and reconstruction examples come from separate recordings.
Fig. 2: MonoEgo workflow. Camera and wrist-constellation calibration precede monocular recording. Offline MonoTag reconstruction combines natural features, workstation anchors, and wrist observations to estimate metric camera and wrist-constellation trajectories; conditional hand estimates and frozen RGB frames form the exported egocentric data.
Method
Odin-01
Odin-02
Odin-03
Metric trajectories: SE(3) alignment
Marker-only
8.577 / 16.2
8.065 / 20.2
8.380 / 30.3
ORB + loose marker
12.271 / 86.3
0.800 / 89.0
0.250 / 49.8
UcoSLAM †
0.055 / 2.8
0.007 / 2.2
0.217 / 8.2
TagSLAM
FAIL
FAIL
0.034 / 19.3
MonoTag (Ours)
0.302 / 78.1
0.196 / 99.5
0.150 / 99.2
TABLE I: Camera trajectory results on Odin recordings (ATE in m / primary-map coverage in %).
Fig. 3: MonoTag trajectories and sparse ORB map points after SE(3) alignment to Odin, without scale fitting. Symbols mark ArUco centres, revisits, scale corrections above 3%, map merges, and the Odin-01 gap and later submap. Only the largest connected map is shown.
Fig. 4: Odin-03 trajectories with and without interval re-anchoring after SE(3) alignment to the Odin reference.
Fig. 5: Stationary wrist-constellation precision with a moving camera. Curves show three-dimensional displacement from each side’s valid sequence mean; missing poses remain gaps.
Fig. 6: Egocentric outputs and offline recovery. (a, b) Annotated RGB with ORB features, marker observations, and hand overlays when available. (c) EGO-04 poses recovered retrospectively in metric units (green); dotted intervals remain unlocalized.
Monocular egocentric human pose estimation is essential for ubiquitous activity monitoring. However, understanding the user's absolute location within the environment remains a challenge. Existing methods primarily focus on relative motion from an initial position, and tend not to account for the wearer's absolute location within an environment. Furthermore, inherent scale ambiguity in monocular vision leads to severe translational drift, limiting long-term tracking without specialized multi-sensor hardware. To address this, we propose MapMonoEgo, a novel framework achieving globally consistent human pose estimation solely from a monocular camera by leveraging a pre-scanned 3D point cloud. We also introduce AIST-Living dataset, a new dataset pairing egocentric video with ground-truth motion in a scanned environment. Experiments demonstrate that our approach significantly outperforms the state-of-the-art baseline, proving its utility for practical monitoring tasks without specialized hardware.
Hiroyuki Deguchi, Ryosuke Hori, Kotaro Amaya +3
1Keio University · 2National Institute of Advanced Industrial Science and Technology
Learning manipulation from human video requires high-fidelity hand-motion reconstruction in metric units. Today's metric hand labels come from studio rigs and instrumented headsets, and both are confined in the same two ways: neither leaves a prepared setting, and neither is checked against an independent reference. Unconstrained head-worn recording promises the opposite trade-off, scaling with the number of people wearing a device. We therefore introduce MEgoVista, an offline pipeline that turns a single unprepared MEgo View recording into metric two-hand and head motion in one gravity-aligned world frame. Three properties set it apart from existing egocentric reconstruction systems: first, it reconstructs in settings studio volumes and tabletop rigs cannot reach, settling hand ownership at detection so bystander hands stay out of the wearer's trajectory; second, it takes its metric gauge from calibrated stereo rather than a monocular prior, installing scale at initialisation so policies receive physical units, not arbitrary coordinates; third, both outputs are scored inside a motion-capture volume against independent Chingmu optical capture, under a protocol that audits its own reference and charges what a method declines to predict. MEgoVista is offered as a measured route from egocentric video to metric hand supervision, one that widens where such labels can be gathered.
Jiangong Xiao, Zhihao Zhang, Yifei Dong +7
Northwestern Polytechnical University · Maniformer · Xi'an Jiaotong University +1
Active scene reconstruction enables robots/UAVs to autonomously plan trajectories and reconstruct environments without costly manual data acquisition. Unlike passive methods, active reconstruction requires real-time construction of high-confidence occupancy maps for collision-free navigation. Existing approaches rely on depth sensors for occupancy map updates, increasing platform cost and weight. To advance spatial intelligence, we aim for a vision-only monocular solution. However, current monocular scene reconstruction methods operate offline and fail to deliver globally consistent dense depth at the frame rates required for robots/UAVs navigation. To bridge this gap, we introduce ActMVS, the first framework for monocular active reconstruction. Our framework integrates a view factor graph construction for informed Multi-View Stereo depth prediction, along with a global depth optimization, to enable the online generation of high-quality, globally consistent dense depth maps. This enables monocular robots/UAVs to maintain reliable occupancy maps for safe trajectory planning during reconstruction. Experiments on Replica datasets demonstrate performance competitive with RGB-D methods. Our code and data are available at https://github.com/TrickyGo/ActMVS.
Guo Pu, Yixuan Han, Zhouhui Lian
Wangxuan Institute of Computer Technology, Peking University