Despite its promise for scaling robot learning, egocentric manipulation data is still scarce today. Collection at scale requires vertically integrating ergonomic hardware with centimeter-precise 3D algorithms, at a precision that has not been publicly demonstrated. To address this gap, we introduce RoboCap, a 250,g six-camera dual-IMU hat designed for in-the-wild egocentric data capture, and the Grounded API, a suite of device-agnostic 3D algorithms tuned for RoboCap. In this report, we demonstrate how hardware, calibration, and 3D algorithms interact to achieve state-of-the-art performance on the public benchmarks: our SLAM across diverse settings and rigs, our depth estimation on egocentric settings, and our hand tracking when adapted to third-party devices.
Figures & tables
Figure 1 : RoboCap camera coverage. Left: azimuth coverage of the front stereo and lateral cameras. Center: azimuth coverage of the eye stereo cameras, visualized forward for legibility. Right: elevation coverage of the front and eye pairs. FoV-overlap in degrees at infinity.
Cameras (6, hardware-synchronized)
Sensor
Global-shutter RGB, 10-bit, 3.0 µm px
Frame resolution
1920×1080
Frame rate
30 Hz
Lens
Kannala-Brandt 4
Focal length
f=590 – 630 px
Exposure
Per-camera AE, bounded ≤10 ms
Table 1: RoboCap sensor suite. The two stereo pairs use fisheye optics, the lateral cameras a wide lens. Stereo overlap is the binocular field at infinity from the calibrated extrinsics. IMU noise parameters are fleet-level estimates from per-device continuous-time calibration [ 16 ] .
Figure 2 : Epipolar residuals before (top) and after (bottom) self-calibration. Identical features are marked in the left and right cameras with a black dot, and epipolar lines are drawn across each pair. Self-calibration eliminates almost all epipolar drift and improves 3D estimation.
Figure 3 : Histograms of the epipolar drift in the eye pair (left), front pair (middle), and front-eye pairs (right), over a subset of 150 deployed devices. We find that rotational correction (Section 4.1 ) is sufficient to restore epipolar drift, which follows a Gaussian centered around 0 ∘ .
Figure 4 : Scene reconstruction over 4 RoboCap sessions. For each session, we generate point clouds (Section 4.3 ) and project them into the SLAM world frame, leveled with the estimated gravity direction (Section 4.2 ). Then, we solve multi-session frame alignment with RANSAC on 3D-projected SIFT features.
Method
Inputs
Short
Medium
Long
Low light
Moving
Avg.
Aria’s SLAM (r) ‡
bino, IMU
90.7
78.5
70.9
84.2
55.0
75.9
ORB-SLAM3 (r) [ 7 ]
mono, IMU
23.0
10.9
11.2
3.1
2.0
10.0
OKVIS2 (r) [ 30 ]
bino, IMU
20.0
11.6
2.6
14.5
4.7
10.7
OpenVINS+Maplab (r)
bino, IMU
27.7
23.4
12.8
19.8
13.9
19.5
Best of other submissions
mono, IMU
34.2
29.3
16.8
25.4
15.5
24.2
Ours
bino, IMU
80.2
61.6
59.9
67.7
41.9
62.3
Table 2: Official LaMAria test-set results (control-point score, higher is better). (r): reproduced by the benchmark authors. “Best of other submissions” is the highest score per category over all other public entries. ‡ Shown in gray as a reference and excluded from the ranking: Aria’s SLAM is the device manufacturer’s own system, developed together with the glasses, and is the solution used to initialize the benchmark’s ground-truth pipeline (its pose recall is therefore not reported). Bold marks the best result among the remaining methods.
Figure 7
Figure 6 : Examples from our hand tracking model, which fuses multi-view images into a single hand mesh. Our model’s pretraining makes it robust to blur, visual ambiguity, complexity, gloves, occlusion, and lighting. See Figures 10 and 11 for more difficult examples.
Scratch
Pretrained
HOT3D [ 3 ]
17.41
14.47
UmeTrack [ 24 ]
19.00
11.49
SHOW3D [ 45 ]
19.34
13.68
Table 5 : Best validation MPJPE (mm) on held-out subjects across datasets, varying only initialization. More details in Appendix A .
POEMv2 [ 64 ]
Ours
DexYCB-Mv
6.69
5.22
OakInk-Mv
8.34
7.25
HO3D-Mv
7.72
13.01
Table 6 : MPJPE (mm) on third-person multi-view benchmarks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10 : Robustness to degraded imaging and extreme viewpoints. Each panel shows the four camera views of a single frame with the predicted mesh overlaid. The model remains multi-view consistent when the hand is blurred by fast motion, when the lens is fogged, when the hand is close enough to the camera to fill the frame, and when the hand is seen through a transparent glove.
Figure 11 : Robustness to hand-object interaction and unusual configurations. Each panel shows the four camera views of a single frame with the predicted mesh overlaid. The model recovers plausible poses while grasping small objects, while holding a tool with the fingers largely occluded by it, during fine manipulation such as drawing, and when the hands appear alongside other body parts.
Figure 12 : Comparison of random vs pretrained weights on egocentric benchmarks. Within each panel the columns are the left and right camera views of a single frame, the top row is the model trained from scratch and the bottom row is the model initialized from our pretrained weights; both are trained until they achieve their best multi-dataset validation metrics. The pretrained model aligns more closely with the observed hand, particularly at the finger joints, where the scratch model tends to mispredict articulation while keeping the global hand position roughly correct.
Held-out (ours)
Official test
subjects
share
subjects
share
UmeTrack
4
10.1%
4
—
SHOW3D
6
20.0%
6
21.0%
HOT3D
2
16.0%
2
—
Appendix
Table 7: Composition of the surrogate held-out sets. UmeTrack: 103 sequences, 12.6% hand-hand interaction against the official split’s 11.6%; real sequences only. SHOW3D: 337 scenes covering 23 of 28 object classes against the official split’s 23; novel-activity rate 12.7% against 12.2%. HOT3D: two held-out subjects per device, matching the official pose track.
views
train
test
split
DexYCB-Mv [ 10 ]
8
25,387
4,951
S0
OakInk-Mv [ 60 ]
4
58,692
19,909
SP2
HO3D-Mv [ 22 ]
5
9,087
2,706
7 seq.
Appendix
Table 8: Multi-view benchmark splits, in multi-view frames, following [ 64 ] . HO3D-Mv comprises the seven sequences: ABF1, BB1, GSF1, MDF1 and SiBF1 for training, GPMF1 and SB1 for evaluation.